Capability Gap

All topics

What better models do to each layer

Three different things get called 'AI will eat this': models doing the crowd's work, models doing the middleman's work, and models creating the budget in the first place. They point in opposite directions and every vertical in the atlas sits in a different one.

medium confidence9 minupdated 2026-08-29ai · automation · synthetic data · forces · labour

"AI will eat this business" is three separate claims wearing one coat, and they have opposite signs.

The first is that models do the labour — the annotation, the sourcing, the first-draft creative the crowd is paid for. The second is that models do the coordination — matching, QC, project management, which is the middleman's own job and the only thing the middleman is paid for. The third is that models create the budget: RL environments, evaluations and expert reasoning traces are line items that did not exist in 2022 and exist now only because models got good enough to need them.

A vertical can sit in all three at once. Expert data is eaten at the bottom, threatened in the middle and inflated at the top, simultaneously, inside one company's revenue line. Treating that as one question is why most forecasts here are wrong.

So what

The score in this atlas asks only the net question: does a better model grow this budget or delete it. But the operator has to answer all three separately, because the mitigations are different. You escape labour substitution by moving up the skill curve. You escape coordination substitution by owning something that is not coordination. You cannot manufacture budget creation at all — you can only be standing in the right vertical when it happens.

Effect one: models eat the labour

This one has already happened, and it is measurable in the accounts of the companies it happened to.

Gilardi et al. showed ChatGPT outperforming crowd workers on text annotation at roughly a twentieth of the cost — in 2023 (PNAS). What the accounts look like three years later: Appen Global fell 21.1% to $127.9M in FY2025 (Appen FY2025 Annual Report); Sama issued redundancy notices to 1,108 Nairobi workers in April 2026 after Meta terminated one contract (TechCabal); xAI cut 500 of ~1,500 generalist annotators in September 2025 while pledging to grow its specialist tutor team tenfold (TechCrunch).

The xAI split is the clearest statement of the mechanism anyone has made: the same buyer, in the same month, cut the cheap tier by a third and multiplied the expensive tier by ten.

So the substitution is real and it is specific to the commodity tier. What it did not do is shrink the sector: the wage ladder in expert data now spans from sub-cent-per-task Remotasks work to $300+/hour for MDs and PhDs on Handshake (AOL/Business Insider). Models compressed the bottom of that ladder and stretched the top, because the reason to pay $300/hour is precisely that the output cannot be synthesised — it is proprietary professional process knowledge that never existed in public text. Brendan Foody's version: labs cannot buy the workflows they want to automate because "their customers don't want to give them data to automate large portions of their value chains, so they need to hire contractors" (TechCrunch).

Effect two: models eat the coordination

This is the underrated one, because it is the effect that hits the operator rather than the crowd.

The middleman's function is matching, vetting, routing, QC and project management. All five are text-and-judgement tasks of exactly the kind models are getting good at, and the sector is already automating them — Mercor built its business on a 20-minute AI interviewer, and OpenAI's Project Mercury hired 100+ ex-bankers directly using its own AI chatbot screen and modelling tests, at $150/hour (Entrepreneur, citing Bloomberg). A lab that can run the funnel itself does not need the funnel operator.

But the evidence cuts hard the other way in the one vertical where the coordination layer has been measured, and it is worth sitting with the shape of it.

Greenhouse data across 640M+ applications shows 244 applications per open role in 2025, up from 116 in 2022, and time to fill up 37% — 43.6 to 59.7 days — even as monthly hires per recruiter rose 122% (TheHireHub). Models made applying free. Making applying free destroyed the signal that inbound applications carried. Destroying the signal raised the price of human filtering, which is why a 20–25% placement fee still clears in contingency recruiting in 2026.

So what

The general form: when a model makes the supply of an input free, the scarce thing becomes discrimination between inputs — and discrimination is the middleman's job. Coordination substitution and coordination inflation are the same mechanism running at different points in the funnel.

That defence does not generalise. It holds where the volume of candidate items explodes and verification is expensive: recruiting shortlists, worker identity, content QC. It fails where the coordination was always cheap — a clipping campaign on a 10% marketplace fee has no judgement in it to protect.

Effect three: models create the budget

The clean cases are new line items with no predecessor.

Anthropic discussed spending over $1 billion annually on RL environments as of September 2025 (The Information, via Epoch AI). Nobody had an RL-environments budget in 2023. Adversarial evaluation is compulsory-by-norm before a frontier launch and had no buyer three years ago. Expert reasoning traces exist because reasoning models made step-level human demonstration valuable in a way that next-token pretraining did not.

Epoch's pricing for that new budget: six-to-seven figures per quarter per contract, $200–$2,000 per task, ~$20K for a website replica, $300K+ for a high-fidelity Slack clone, and exclusive deals at 4–5x non-exclusive (Epoch AI). The exclusivity multiple is the tell: labs are paying largely to deny the asset to rivals, which is not a price anyone pays for a commodity input.

The counterweight, from the same source: Anthropic's $1B on environments sits against OpenAI's **$19B 2026 R&D compute budget** (Epoch AI). Budget creation is real and it is an order of magnitude smaller than the compute line it accompanies.

Which vertical sits where

VerticalLabour eaten?Coordination eaten?Budget created?Net (ai score)
Expert data for frontier labsBottom tier gone; expert tier growingPartly — labs run their own funnels (Project Mercury)Yes — environments, evals, traces4
Adversarial evals and red-team crowdsNo — models are the target, not the substituteNo — the crowd self-organises around leaderboardsEntirely new budget5
Robotics teleoperation and physical-world dataNo — the bottleneck is physical demonstrationsWeaklyYes — physical AI took $47.4B across 521 deals in H1 2026 (Crunchbase)4
Contingency recruiting marketplacesSourcing partly automatedInverted — signal destruction raised filtering valueNo3
Expert networksNo — the product is indemnified accessPartly — matching is automatableUnproven that labs buy from them at all3
Paid creators, clipping and UGC ad opsYes — synthetic video runs $5–$25 per finished asset (SparkUGC) [WEAK]Yes — the take was 9–18% for routing, which is automatableNo3
Voice, speech and low-resource language dataLargely — commodity speech collection has collapsed in priceYesOnly in long-tail languages and consent provenance2

The pattern: the verticals that score well are the ones whose constraint is not textual. Red-teaming needs someone who wants to break things. Robotics data needs a rig and a floor. Expert data at the top of the ladder needs someone who actually did the job at Goldman. Everything whose constraint was "arrange text at scale" is on the wrong side.

The synthetic-data argument, both sides properly stated

The strongest case for displacement. The commodity tier is already gone, demonstrated rather than predicted — Gilardi's 1/20th cost result, Appen Global at −21.1%, Sama's 1,108, xAI's 500. Contrary lists synthetic data and RLVR displacement as a top-tier risk to Mercor specifically: if verifiable rewards substitute for human feedback, "the entire AI training data market" contracts (Contrary). And the buyer's own stated intent runs this way — ICONIQ finds talent's share of cost falls while inference rises as AI products move from beta to scale, with internal AI infrastructure spend going 11% of revenue in 2025 → 16% in 2026 → 19% in 2027 (ICONIQ, State of AI 2026). 78% of AI companies are rethinking staffing, 45% planning a different role mix, 33% planning smaller teams (same source). The cost structure is being deliberately rotated away from people.

The strongest case against. If synthetic data were displacing humans, the revenue curves would have bent by now and they have not: Mercor booked $614M gross in H1 2026, +70% YoY (The Information, via BigGo); Handshake went zero to $1B gross in about fifteen months (Dealroom); micro1 5x'd in eight months (Dataconomy); Innodata is guiding to 40%+ growth on audited numbers (StockTitan). Kang's estimate is that 2024 data-labelling spend ran at **3.1x the marginal compute cost** of frontier training runs, having grown 88x from 2023 to 2024 while compute cost grew 1.3x (ddkang) — the $100/sample assumption does a lot of work there, so treat it as directional. And Epoch's operational finding is that human expertise, not technical capability, is the bottleneck in producing environments (Epoch AI).

Neither side is settled by the research. RLVR work in 2026 is genuinely contested — one line argues it "makes models faster, not smarter" (Promptfoo) [WEAK], while active-label-acquisition papers find verifiable rewards still need human labels where the model's self-belief misleads it (arXiv).

Read

The market is not being eaten, it is bifurcating — and the bifurcation is fast enough that a vendor can be on the wrong side of it within two years. Scale and Appen are what the losing side looks like in the accounts. The winning side buys that growth at a 27–33% gross margin with ~91% of revenue from foundation-model companies, which is a different problem entirely.

Where the record is thin

Gap in the record

Nobody has published a lab's human-data budget. The two questions that matter most — how much of the environments budget is genuinely incremental rather than reallocated from annotation, and whether the expert tier's margin holds as models close on professional reasoning — have no source at all. The ICONIQ cost-rotation finding is stated intent in a survey, not observed spend. And 244 applications per role is measured, while "human filtering therefore gets more valuable" is an inference from it, not a measurement of what buyers paid.