Capability Gap

All verticals

Expert data for frontier labs

Labs buy throughput of credentialled human labour — annotation, preference data, reasoning traces, RL environments — and pay nine figures for it at a staffing margin.

crowdedhigh confidence12 minupdated 2026-08-29ai labs · data · expert labour · rl environments
Who is buying
Frontier AI labs — OpenAI, Anthropic, Google, Meta, xAI, Microsoft
Who is selling
Doctors, bankers, lawyers, PhDs, competitive programmers, senior engineers
The spread
27–33% gross margin (Mercor, leaked); 42.8% published take (Prolific)
Size of the pool
~$8.5B/yr gross across 50+ vendors; probably $3–4B net
Read
The budget is real, urgent and growing. The margin is a staffing margin, the buyers are two companies wearing five names, and every seat at the top table is taken.
Budget depth
5
Supply difficulty
3
Spread
2
Holdability
2
AI direction
4
Speed to first dollar
4

A frontier lab cannot buy the thing it most wants. It wants the workflow inside a restructuring banker's head, the diagnostic reasoning of an oncologist, the way a litigator actually drafts. None of that is on the open internet and none of the companies that own it will hand it over. So the labs do the only legal thing available: they hire the humans, one hour at a time, through a broker.

That broker market ran to roughly $8.5B of gross billings in the year to mid-2026, with the top four — Scale, Surge, Mercor, Handshake — taking more than 75% of it across 50+ vendors (Pebblous, citing Deedy Das's vendor map) [WEAK — analyst synthesis]. A bottom-up cross-check of the named companies lands at $7–9B — but that cross-check adds Mercor's gross to Appen's net, so it corroborates the total by luck rather than by construction; see What we could not establish, item 8. But almost all of it is gross payment volume. Contractors keep 60–70%. Net revenue to the whole vertical is probably $3–4B — see GMV is not revenue before you compare any figure here to a software company.

The most surprising fact in the market is not the growth. It is the margin. Internal Mercor documents obtained by The Information in July 2026 put the fastest-growing company in the sector at a 27% gross margin in 2025, rising to 33% in Q2 2026, with labour absorbing about two-thirds of revenue (via BigGo Finance). That is a staffing margin. Robert Half's contract-staffing segment runs 39%. The best-executed services business on earth, Accenture, runs 32% and trades at 1.56x revenue. Mercor was in talks at $20B.

Whose budget

Score: budget 5. This is the deepest, most urgent buyer pool in the atlas, and the only reason the rest of the scores are worth arguing about.

No lab has ever disclosed a human-data budget. Every number below is a leak, an estimate or a third-party reconstruction.

AnchorFigureSource
Anthropic, RL environments alone, forward 12 monthsdiscussed >$1BTechCrunch, Epoch AI
Google spend with Scale~$150M (2024), ~$200M planned (2025)Sacra
Major labs, human data, each~$1B/yr [WEAK]HeroHunt
Top six Chinese labs, spend at US labellers~$500M/yr [WEAK]AI Weekly summarising Forbes
Disclosed contract sizes"eight- or nine-figure" (Edwin Chen)Inc via Yahoo

Chen's remark is the only concrete contract-size disclosure anywhere in the vertical. Everything else is inference.

Two counterweights. First, Epoch's framing: Anthropic's $1B on environments is modest against OpenAI's projected $19B of 2026 R&D compute (Epoch AI). The human-data line is not on a path to rival the compute line. Second, and cutting the other way, Daniel Kang's analysis estimates 2024 data-labelling spend at **3.1x the marginal compute cost** of frontier training runs, with labelling revenue growing 88x from 2023 to 2024 against 1.3x for compute (ddkang). His $100/sample assumption does a lot of work; treat it as directional.

Can you get the supply

Score: supply 3. Hard, but demonstrably buyable — and demonstrably shared.

The supply is credentialled professionals with day jobs: ex-Goldman and McKinsey staff, Latham & Watkins lawyers, Mount Sinai clinicians (TIME). Assembling it means running an identity-verified funnel, paying across 45+ countries, and carrying contractor liability. That is genuinely hard, and it is the vendors' actual product — see The law is about to arrive.

But three facts cap the score. Handshake stood up a $1B gross business in fifteen months because it already owned the network: 17–20M students, 1,600+ institutions, ~500,000 PhDs in-network (Lenny's Newsletter). Scale and Mercor had been recruiting off Handshake; Handshake simply stopped selling the picks and shovels. Second, the supply is non-exclusive — the same annotators work across Surge, Outlier and Mercor sub-brands, so nobody's pool is anybody's moat. Third, OpenAI hired 100+ ex-bankers directly at $150/hour for Project Mercury, complete with its own 20-minute AI interview and weekly Excel model (Entrepreneur, citing Bloomberg). A lab built a Mercor internally, screening funnel and all.

Headcount claims should be read sceptically. Mercor reports 30,000+ active vetted contractors (Contrary) while Fortune cites a network of "five million domain experts" (Fortune). Surge shows 50,000 (Sacra) against Wikipedia's ~1 million (Wikipedia). These are almost certainly active-versus-registered, but no source says so.

What the spread looks like

Score: spread 2. Mercor's leaked 27→33% is the best evidence in the sector, and it is a services margin.

Three numbers bracket the truth. Prolific publishes its take: "our platform fee is usually 42.8% for corporate customers, and a discounted 33.3% for academic or non-profit customers" (Prolific pricing) — the only public take rate in the vertical, charged on top of participant pay. Mercor's leaked gross margin is 27–33%. Handshake and Micro1 both pay contractors 60–70% of gross, implying a 30–40% take (Dealroom; Dataconomy). See What a rake can actually be.

The public comparables are unkind. Appen ran a 40.3% gross margin in FY2025 (Appen FY2025 Annual Report); Innodata reported 49% adjusted gross margin in Q2 2026 (StockTitan). Both low-glamour outsourcers out-margin the $10B unicorn.

So what

Two escape routes from the labour margin exist, and both are visible in the data. Micro1 claims 80–90% gross margin on off-the-shelf datasets resold to multiple clients (Dataconomy). And Epoch reports that exclusive RL-environment deals price at roughly 4–5x non-exclusive (Epoch AI) — meaning the labs are paying mainly for denial, not for data. Sell the same asset many times, or sell the promise not to. Selling hours pays 30%.

Epoch's other pricing anchors: contracts of six to seven figures per quarter, ~$20,000 for a website replica, ~$300,000 for a Slack-grade product clone, and $200–$2,000 per individual task. Mercor's own APEX benchmark cost >$500K for 200 tasks — about $2,500 each (TIME).

Can you hold it

Score: hold 2. The buyers are two companies wearing five names, and they leave in days.

~91% of Mercor's H1 2026 revenue came from AI foundation-model companies, dominated by OpenAI and Anthropic (The Information via BigGo). Innodata's top two customers are 71% of Q2 revenue (StockTitan). Appen's top five are 74.3% (Appen AR2025). This is One customer is a binary event in its purest form.

Worse, the buyers compete with each other, which makes neutrality the product. Meta paid $14.3B for a 49% non-voting stake in Scale in June 2025 and took its CEO (Wikipedia; Computerworld). Within weeks Google cut ties, OpenAI departed, Microsoft pulled back and xAI began exploring alternatives (Sacra). By July 2025 Scale had cut 200 FTEs and ~500 contractors. By August 2025 Meta itself was routing work to Surge and Mercor (TechCrunch). The largest financing in the category's history was simultaneously its largest customer-loss event.

Nothing holds the supply either. Contractors work across platforms; the frontier of what a vendor can charge is set by the buyer's willingness to hire directly, and Project Mercury proved they will. What the vendors genuinely own is unglamorous: payroll compliance in 45 countries, contractor liability, and the ability to field 5,000 vetted doctors in three weeks. Mercor's numbers price that at ~30% of gross, not 80%.

What AI does to it

Score: ai 4. Models create this budget at the top and are eating it at the bottom.

The market is not being eaten; it is bifurcating. Gilardi et al. showed ChatGPT already outperformed crowd workers on text annotation at roughly 1/20th the cost — in 2023 (PNAS). The commodity tier is largely gone, and you can see it in the accounts: Appen Global fell 21.1% to $127.9M in FY2025, and Sama issued redundancy notices to 1,108 Nairobi employees in April 2026 after Meta terminated its contract (TechCabal).

At the same time xAI cut 500 of ~1,500 generalist annotators in September 2025 while pledging to grow its specialist tutor team tenfold (TechCrunch). That single decision is the whole What better models do to each layer in miniature: the same budget, moved up a tier.

The bear case is not settled. Contrary lists synthetic data and RLVR displacement as a top-tier risk to Mercor — if verifiable rewards substitute for human feedback, "the entire AI training data market" contracts (Contrary). One 2026 line of work argues RLVR "makes models faster, not smarter" (Promptfoo) [WEAK]; active-label-acquisition papers find verifiable rewards still need human labels where the model's self-belief misleads it (arXiv). Against that: Mercor doubled gross revenue in four months, Handshake went 0→$1B in ~15 months, Micro1 5x'd in eight, and Innodata is guiding to 40%+ growth on audited numbers. If synthetic data were displacing humans, those curves would have bent. They have not.

What would kill it

What would kill it

Four things, in rough order of probability.

One buyer decision. Appen is the base rate: peak market cap ~$4.3B in 2020 with 80% of revenue in five clients, Google terminated in January 2024, and the shares fell 40–41% in a day. The Verge computes the drawdown at 97% from peak. Nothing about the current cohort's concentration is milder — it is worse.

Reclassification. Suits are pending against Mercor, Surge, Scale and Handshake. Contrary's assessment is that reclassifying contractors as employees would make the ~35% take "untenable" once benefits, overtime and compliance land (Contrary). This is a margin question, not a legal footnote.

Fraud eating the premium. The buyer is paying for "verified human expert" and the verification is leaky. Inc obtained internal logs showing Scale's Google programme was plagued by spam for 11 months — contributors using ChatGPT despite a prohibition, faking advanced degrees, accounts logged 18+ hours daily implying shared credentials, and one incident where "the Allocations Department had dumped 800 spammers into our team" (Inc). Mercor now runs a public Kaggle competition to detect interview cheating (Kaggle). See Who is actually on the other end.

Security and supply reputation. Mercor lost ~4TB in a March/April 2026 breach — passports, SSNs, biometric face and voice data — drawing six class actions and causing Meta to pause work (Staffing Industry Analysts). Meanwhile wages fall at the bottom of every platform: Mercor moved Meta "Musen" workers to "Nova" at $16/hr from $21/hr (Forbes); Outlier baseline reportedly fell from $40–50 to $15–20 (Breaking Even) [WEAK]; Handshake's "Project HH" paid 20–50% of earned amounts in May 2026 (Breaking Even) [WEAK]. Appen's crowd NPS fell from 33 to 22 on audited disclosure. The pools are shared; a reputation collapse on Reddit is a real cost of goods.

Who is already there

Score: speed 4. Standing starts have reached nine figures inside eighteen months — repeatedly.

CompanyRevenue basisLatestValuationRead
Mercorgross$2.0B run rate (Jun 2026)$10B; $20B in talksThe archetype, and the leak
Surge AIgross$1.2B (2024)$15B or $25B — unresolvedBootstrapped, profitable, 130 staff
Scale AImixed~$2B est. 2025; "surpass $1B" 2026$29B (Jun 2025)What losing neutrality costs
Handshake AIgross~$1.0B AI arm (Apr 2026)$3.5BOwned the supply already
micro1gross$500M run rate (Aug 2026)$500M (stale)80–90% GM on resold datasets
Turinggross$300M+ (2024)$2.2BStaffing spread, stale numbers
Invisible Technologiesnet-ish$134M (2024), $15M EBITDA$2B+Discloses real profit
Prolificnet$350M est. (Apr 2026)undisclosedPublishes 42.8%
Appennet$230.8M (FY2025)A$330M mkt capThe base rate
Mechanizenoneundisclosed$500M → $1.5B talksLicence-and-hire

Two structural entrants matter beyond the league table. Mechanize raised $9.1M at ~$500M and was in Google licence-and-hire talks at $1.5B+ (TFN) — the template for selling environments rather than hours. And iMerit, a decade-old annotation company with an expert network, sold to EXL for up to $310M in the same quarter Mercor was talking at $20B (EXL 8-K via StockTitan). The market prices growth rate and lab relationships, not the labour network — which is exactly what What the public market pays for labour should make you nervous about.

Where the record is thin

Gap in the record

No lab has disclosed a human-data budget. Every figure in "Whose budget" is a leak, an estimate or an aggregate.

Gap in the record

Only six take rates are pinned down anywhere: Mercor (27→33%, leaked), Prolific (42.8%, published), Micro1 (30–40%, implied and internally inconsistent in its source), Handshake (60–70% payout), Appen (40.3% audited FY2025) and Innodata (40% on the TTM, 49% adjusted in Q2 2026). Surge, Scale, Turing, Labelbox, Toloka, Snorkel and Uber AI Solutions disclose nothing.

Gap in the record

Scale AI's actual 2025 and 2026 revenue could not be established. Estimates run from ~$870M (2024) to ~$2B (2025E) to "surpass $1B" (2026) on definitions that are certainly incompatible. Surge's 2025 and 2026 revenue is equally unknown — the last credible figure is $1.2B for 2024.

Gap in the record

No net-revenue league table for this vertical exists publicly. Everyone reports GMV. Every multiple you will read, including the ones on this page, is a range.

Not verified

The Chinese-lab revenue figures — ~2% for Mercor, ~$500M of aggregate annual spend at US labellers — are second-hand from Forbes reporting that could not be fetched directly. Treat as unverified.

Two more holes worth naming. Whether OpenAI's Project Mercury contractors were sourced directly or through a vendor is not in the reporting, and it materially changes what that case proves about in-housing. And the active-versus-registered contractor discrepancies — 30,000 vs 5 million at Mercor, 50,000 vs 1 million at Surge — are unresolved by any source.