Capability Gap

How to evaluate one of these

Six scores, one to five, higher always better for the operator — what each is trying to capture, where each misleads, and the three disqualifiers that override any total.

high confidence11 minupdated 2026-08-29framework · scoring · rubric · diligence · method

Every vertical page in this atlas carries six scores. They exist so two markets can be compared in one line, and they are dangerous for exactly that reason. Read this page before using the compare table: a total is a blunt instrument, and the three disqualifiers at the bottom are not.

budget — how deep and urgent the buyer's money is

5 = a lab that will sign a nine-figure contract this quarter. 1 = a seed startup paying by card.

ScoreAnchorWhy
5Expert data for frontier labsSix-to-seven figures per quarter is normal; Surge cites nine-figure contracts
4Expert networks~$3B category, ~11,200 client firms, paying with someone else's money
3Contingency recruiting marketplacesReal budgets, but a fee measured against one salary
2Design, video and content production$50k–$300k a year across creative and content combined
1Nothing in the atlas scores 1; the anchor marks the floor

This is not market size. It is whether the money is already approved, already urgent, and larger than the friction of buying. How it misleads: budget measures the buyer's cheque, not your access to it, and depth correlates with concentration — Expert data for frontier labs runs budget 5, hold 2, because the counterparties that make the budget enormous also make it a coin flip.

supply — how hard the seller side is to assemble

Hard is good. 5 = you cannot buy this supply with money alone. 1 = anyone can run the same ads tomorrow.

ScoreAnchorWhy
5Adversarial evals and red-team crowdsElite jailbreakers: small, self-selecting, not recruitable by job ad
4Data-centre labour and site brokerageJourneyman electricians against a 349,000–499,000-worker 2026 shortfall
3Expert data for frontier labsHard — but Handshake assembled it in fifteen months off an owned graph
2Paid creators, clipping and UGC ad opsClippers aged 16–24, unvetted, multi-homed
1Design, video and content productionGlobally abundant, reachable through ordinary channels

How it misleads. "Hard to assemble" is not "hard for a competitor who owns an adjacent asset" — Scale AI and Mercor were recruiting PhD annotators off Handshake until Handshake stopped selling picks and shovels. Supply you had to manufacture is a training-cost sink whose output walks out of the door (Andela). And pool sizes are worthless: Mercor reports 30,000+ active vetted contractors against a Fortune claim of five million. Score the funnel you have watched work.

spread — what fraction of the money passing through you, you keep

5 = 60%+. 1 = under 10%.

ScoreAnchorWhy
5Expert networks$200/hr to the expert, $800–900/hr to the client, four decades
4Outbound and GTM-as-a-service~55–75% gross margin [UNVERIFIED — inferred from two published rate cards]
3Robotics teleoperation and physical-world data40–65% gross on quality teleoperation
2Expert data for frontier labs27–33% gross margin (Mercor, leaked)
1Compute and capacity brokerageLow single digits against a public price index

How it misleads. Take rate is not gross margin: Upwork takes 18.7% and runs a 77.2% gross margin; Robert Half marks contract labour up around 65% and runs 39.0%. A gross-booked staffing spread and a net-booked marketplace fee are not the same asset — see GMV is not revenue. Worse, the highest spreads in this atlas sit on the weakest evidence. Design, video and content production's 70–88% is arithmetic on Glassdoor pay data whose own source gives two bands thirty points apart; Adversarial evals and red-team crowds's >90% is inferred with no revenue disclosure anywhere. A 4 on a guess is not a 4.

hold — whether they can transact around you, and survive without you

5 = they cannot leave. 1 = they meet once and you never see them again.

ScoreAnchorWhy
5Nothing scores 5; the honest ceiling for this model is 4
4Expert networksChaperoning, subscriptions, 22 of GLG's top 25 clients retained over a decade
3Forward-deployed engineeringDelivery risk and embedded deployment, against a project-shaped budget
2Expert data for frontier labs~91% of Mercor's H1 2026 revenue from foundation-model companies
1Paid creators, clipping and UGC ad opsA brand and a clipper who have met once transact by direct message

How it misleads. It nets two independent risks: Getting cut out asks whether the two sides can transact around you, One customer is a binary event whether one customer leaving ends you. A vertical can be sticky per-relationship and still be one procurement decision from death — Appen had multi-year embedded contracts and 80% of revenue in five clients. When the two diverge, score them separately and take the lower.

ai — whether capable models grow this budget or delete it

5 = models create the budget. 1 = models are eating the work.

ScoreAnchorWhy
5Adversarial evals and red-team crowdsThe budget line did not exist before capable models did
4Expert data for frontier labsNet of a created budget at the top and a deleted one at the bottom
3Contingency recruiting marketplacesSourcing gets cheaper; the fee is unchanged
2Voice, speech and low-resource language dataSynthesis erodes the collection side faster than demand grows
1Design, video and content productionGenerative tooling is directly substituting for the deliverable

How it misleads. Models act on three layers at once — labour, coordination and budget — and one digit nets them. Expert data for frontier labs at 4 hides the sector's sharpest fact: xAI cut 500 of ~1,500 generalist annotators in September 2025 while pledging to grow its specialist tutor team tenfold (TechCrunch). Same budget, one tier up. If you sit in the tier being vacated, a 4 is a 1. See What better models do to each layer.

speed — time from standing start to a real invoice

5 = weeks. 1 = years.

ScoreAnchorWhy
5Outbound and GTM-as-a-serviceA retainer closes in weeks against an empty pipeline
4Expert data for frontier labsHandshake, Mercor and micro1 reached nine figures inside 18 months
3Forward-deployed engineeringEnterprise deployment cycles
2Data-centre labour and site brokerageTrades licensing, site access and build schedules
1Anything needing a regulatory approval before the first invoice

How it misleads, more than any other score: speed runs inversely to everything defensible. The three verticals scoring 5 are Paid creators, clipping and UGC ad ops, Outbound and GTM-as-a-service and Design, video and content productionwatch, crowded, avoid. High speed beside low supply describes a business anyone can start next week, including your customer.

The three disqualifiers

These override any total. A vertical can score 24 out of 30 and be uninvestable on one of them.

What would kill it

1. Buyer concentration above 25%. Above ~15% of revenue a single customer is a binary event, not a growth line; above 25% you are a division of that customer with a different cap table. TaskUs discloses Meta at 26%. Appen had 80% in five clients at a US$4.3B peak, lost Google in January 2024, and is down 97%. Mercor is at ~91% from foundation-model companies. And where buyers compete, neutrality is the product: Meta's $14.3B for 49% of Scale AI was simultaneously the largest financing and the largest customer-loss event in the category's history.

2. A supply anyone can assemble with an ad budget. If the answer to "how did you get the supply" is "we advertised," your competitor's answer is the same and their price is lower — and the buyer already holds a cheaper comparator: offshore headcount rose from 24% to 30% of allocation year on year at a 40–50% cost arbitrage (SaaStr on ICONIQ).

3. A transaction the two sides can repeat without you. One-shot, high-ticket, low-friction is the maximum-leakage corner and no clause fixes it. Twenty years in, with full telemetry, Upwork's 10-K still calls circumvention losses "difficult or impossible to measure."

How to use it

Answer in this order — the cheap questions are also the fatal ones.

  1. The three disqualifiers. Name your five largest prospective buyers and ask whether losing the largest ends the company; try to source ten suppliers with an ad; describe the second transaction between a buyer and a supplier you introduced.
  2. budget. Name three buyers and the line item the money comes out of. If it is a headcount line, you are priced against a salary they can look up.
  3. supply. Spend a small fixed sum on the sourcing channel and count qualified applicants in one week. That number is your score.
  4. hold. Ask whether anything you do gets harder for the buyer to replicate as they get more sophisticated. If nothing does, you have speed bumps.
  5. spread. Get one real quote for the buyer's alternative and one real supply rate. The spread is the difference, after rework and payout costs.
  6. ai. Attempt the task yourself with a current model, honestly, and time it.
  7. speed. Try to get one invoice paid. See The first ninety days.

Worked example: scoring data-centre labour from scratch

budget 5. Developers and hyperscaler build programmes running a delay clock reported at $14.2M per month on a 60MW project, with projects that once peaked at ~750 workers now needing 4,000–5,000.

supply 4. Journeyman electricians against a shortfall of 349,000–499,000 additional workers in 2026 alone and a ~30% wage premium. Not 5: the supply is licensed and legible, so a capitalised staffing incumbent can chase it.

spread 3. Conventional construction staffing runs 30–40% markups and no AI-native intermediary has published anything — the 3 is incumbent comparables, not an observed vendor.

hold 3. Site access and multi-project relationships create friction, but a developer who has met a local's business manager can hire directly next build.

ai 4. Models do not do this work and the buildout exists because of them. Not 5: the budget is downstream of compute demand, not created by the model.

speed 2. Licensing, site access, build schedules.

Total 21 of 30, near the top of the atlas. Now the disqualifiers. Concentration is a live failure — a handful of build programmes, and one deferral would be an Appen-shaped event. Supply passes. Repeat-transaction risk is real and unpriced.

Not verified

Then the honest discount. The research behind this vertical has demand evidence and no vendor evidence at all — every figure describes the buyer or generic construction staffing, and the $14.2M/month delay cost is single-source and single-project-size. A 21 built on one side of the market is not a 21. This is what a properly applied framework looks like when the answer is "the score is a range, and the range is wide."

So the atlas takes the discount rather than describing it. Data-centre labour and site brokerage carries spread 2, not the 3 above, on two grounds the from-scratch pass glossed: a 30–40% markup is a 23–29% margin, and the band is borrowed from a different buyer in a different market. Total 20, and the stance is watch rather than build, because a vertical with no observed seller is a vertical to go and measure. Compare Robotics teleoperation and physical-world data, which carried an identical six until this pass and has real prices on both sides of its trade — two verticals should not share a vector when one of them has half the evidence.

Scores are judgements

A total is a blunt instrument. The six are neither independent nor equally weighted: speed runs negatively against supply across the whole atlas, budget negatively against hold. Summing correlated judgements produces a number that looks like measurement and is not.

An unargued score is worse than no score. Every vertical page justifies its six in the body. Where a justification rests on a [WEAK] or [UNVERIFIED] source, the score is a placeholder — spread and ai are where this bites hardest, because both are usually inferred rather than observed.

Re-score rather than trust these. These come from one evidence base at one date, with the gaps named on every page. One afternoon of primary conversations — a lab procurement contact, a contributor cohort, one real rate card — beats all of them. What is durable here is the order of the questions and the three disqualifiers. The digits are an invitation to argue.