Why Frontier AI’s Costs Keep Compounding
The economic case for changing architectures
AI | ECONOMICS
Why Frontier AI’s Costs Keep Compounding
The economic case for changing architectures
Photo by Arno Senoner on Unsplash
AI companies are paying experts to write rules in natural language, one case at a time, by the hundreds of thousands of criteria. Informal axiomatization at industrial scale. How much does it cost?
Let’s look at the numbers.
OpenAI’s HealthBench rubric ran to roughly 50,000 evaluation criteria — and that was for a subset of emergency medicine, not all of it. Each criterion was authored by a physician.
The standard expert-data marketplaces pay these physicians by the hour. Surge AI pays medical fellows $250 to $450 an hour, and VC partners and C-suite executives $500 to $1,000.
Mercor — the fastest-growing platform, with 30,000 PhD-and-domain-expert contractors — pays an average of $95 an hour, with STEM specialists at $90–110. Scale AI sits at the lower end of the market with STEM experts at $30–50 an hour and generalists at $15–30. LinkedIn just entered the market, paying specialists up to $150 an hour. The expert-data sector pays out more than $1.5 million per day at Mercor alone.
If we do a conservative back-of-envelope estimate for a single benchmark — two hours of expert time per criterion to draft, review, and revise, at $200 an hour for a medical specialist — a 50,000-criterion rubric costs roughly $20 million.
The actual figures are higher; Surge tops out at $450 an hour for medical fellows, and the iteration cycle adds rounds of revision against model outputs by the same domain experts. A single rubric for a single subfield of one specialty: tens of millions of dollars.
Now do the same exercise for every domain. Every specialty within medicine. Every branch of law. Every subfield of finance, of physics, of chemistry, of engineering. Every edge case in every one of those domains.
The cost to capture the rules of human expertise this way isn’t measured in millions per rubric. It’s measured in billions per domain, across an unbounded number of domains.
Forever and ever
Wouldn’t this be a one-time cost? Pay the experts, build the rubrics, train the models, done.
No. The patterns a model learns from a rubric don’t generalize past the cases the rubric covers.
A rubric for chest-pain triage doesn’t teach the model how to handle stroke triage. A rubric for breach-of-contract analysis doesn’t teach it to evaluate intellectual property disputes. Every new domain requires a new rubric set. Every new scenario inside a domain requires expansion. Every edge case that wasn’t in the original criteria requires the experts to come back and write more.
New model versions add to the cost. Each new release of a frontier model has to be re-evaluated against the rubrics. The rubrics themselves get expanded with new criteria for new capabilities being tested. A rubric written to evaluate last year’s capabilities doesn’t fully cover what a current frontier model can do. The benchmark grows with the model, and the cost of maintaining the benchmark grows with it.
The rubrics serve as both evaluation criteria and training signal. RLHF-as-judge, scoring loops, the next round of fine-tuning data — all of it draws on the same rubrics that were written to grade the previous round. Both uses require the rubrics to keep growing as the models do.
The cost compounds with very new domain, every new scenario, and every new model release adds to the rubric work.
This isn’t capital expenditure that produces an asset. It’s operational expenditure that produces consumables, and the consumables get used up faster than the field progresses.
Two different cost curves
The cost of informal axiomatization scales with the number of cases the system has to handle. The cost of formal axiomatization scales with the number of rules in the system. For any non-trivial domain, the cases vastly outnumber the rules — because rules compose, and cases don’t.
Take tax law. The rules of the U.S. tax code, while large, are finite. The cases the rules apply to are effectively infinite — every taxpayer, every transaction, every fiscal year. Formalize the rules once, and the system applies them to every valid case at no additional cost. Or write rubrics for cases, and pay per case, forever.
Capex asymmetry
Both approaches require capital expenditure. Formalization needs domain experts, knowledge engineers, and engineering effort to capture the rules, the entities, and the bridge logic between layers. The current pure neural architecture needs non-linear scaling of criteria, GPU clusters, data centers, and multi-year compute contracts. These are real upfront cost in both cases.
The difference is what the capex produces. Formalization capex produces an asset whose coverage is determined upfront, and every valid case within that coverage runs at zero marginal cost.
Pure-neural capex produces an asset whose coverage has to be continuously extended — new domains, new scenarios, new edge cases all require new rubrics, new expert labor, new training compute. This pattern is the operational consequence of relying solely on informal, ad hoc assets.
The upfront cost of the alternative — formalization — is real, and a fair worry is whether bounded-domain formalization can actually be built fast enough and cheaply enough to matter. A later piece in this series takes that question up directly. The structural cost asymmetry stands either way.
Data wall as forcing function
Frontier models have consumed essentially all the high-quality text on the public internet. This is the ‘data wall’ — a known problem in the industry, and the reason the rubric activity exists in the first place. If the next gains have to come from somewhere other than scraped text, expert-written rubrics are one of the few sources left.
The same dynamic shows up with experts. There are a finite number of PhDs in any specialized field. A finite number of hours those PhDs have to sell. And the AI companies are competing for the same experts at the same time. Surge, Mercor, Scale, Outlier, LinkedIn — every marketplace is bidding for the same pool, and the pool isn’t growing.
Formalization changes the math on expert time. A formal system for arithmetic generates infinite valid equations from one upfront investment. A formal system for contract law generates infinite valid contracts the same way. The expert hours that go into building the system compound — they produce a structured asset that keeps generating valid reasoning long after the experts have moved on. The expert hours that go into writing rubrics don’t compound. Each rubric covers the cases it was written for and no others.
The cost curve from the previous section would force a pivot eventually. The data wall forces it now.
Market signals
A few years ago, AI companies were paying crowd-sourced workers a few dollars an hour to label images and rate chatbot outputs. Now they’re paying physicians $450 an hour to write conditional rules. They raised their rates because cheap labels stopped producing model improvements.
The premium is evidence about what the labs themselves believe. They’re paying for structured expert reasoning because they’ve concluded they need it. What they’re getting in exchange — natural-language rubrics, one case at a time — is a degraded substitute for formalization. They’re paying for a property of formal systems and receiving prose that imitates the surface of it.
Accelerating timeline
Pressure on the rubric model is building now. Rubric costs go up per criterion, per domain, per model version. Expert marketplaces compete for the same physicians, lawyers, and PhDs while no new supply appears. Frontier model improvement is running out of new training data and new expert hours at the same time.
If the economics force a pivot and the timing forces it soon, why hasn’t the pivot already happened?
That’s the question the next piece in this series takes up.
Part 2 of a series on AI and formalization. Part 0 covered the structural case for formalization. Part 1 covered what the major AI companies are doing instead — paying experts to write rules in natural language for neural models to learn back. The next parts take up why the industry hasn’t pivoted (Part 3), what hybrid neurosymbolic architectures actually look like (Part 4), and the bounded domains where they should get built first (Part 5).
메타데이터
- post_id
- 03bb7decd43f
- slug
- why-frontier-ais-costs-keep-compounding-03bb7decd43f
- url
- https://medium.com/@cfeusier/why-frontier-ais-costs-keep-compounding-03bb7decd43f
- canonical_url
- https://medium.com/@cfeusier/why-frontier-ais-costs-keep-compounding-03bb7decd43f
- author_url
- https://medium.com/@cfeusier
- status
- ok
- fetched_at
- 2026-06-17 15:37:45