This is the most important page in the atlas, and if you only read one page before acting on any of it, read this one.
Everything else here is an argument built on evidence of varying quality. This page is the list of places where the evidence runs out, splits in two, or was never collected by anyone. Some of these are failures of retrieval and could be closed in an afternoon. Some are facts that no one on earth has written down. The difference matters, so it is marked.
Part 1 — The questions that would most change a decision
Ranked by how much the answer would move an operator's judgement, not by how interesting they are. Each carries what it would take to close it and how hard that is.
1. Do expert networks sell to frontier labs at all?
Why it is first. Expert networks hold a 70–80% take that has survived four decades, three compliance scandals and every disintermediation attempt. That is the single most durable rake in the atlas. Meanwhile the labs' actual expert sourcing has gone to Mercor, Handshake and micro1 — companies running 27–33% gross margins. GLG, AlphaSights, Third Bridge and Guidepoint have the supply, the compliance apparatus, the chaperoning infrastructure and the recruiting machine, and appear not to be selling into the fastest-growing buyer of expert hours in the world.
If they are selling and it is simply undisclosed, the incumbents are the strongest position in the sector and nobody has noticed. If they are not, the most valuable unexploited arbitrage in the sweep is sitting in plain sight, and the reason it is unexploited is almost certainly organisational — an incumbent sales force that knows how to sell to PE associates and not how to sell to a research lead.
What we have. Nothing. Directed searching produced no evidence in either direction. What is observable is only adjacent: the networks have built AI-moderated call products, i.e. they are buyers of AI rather than suppliers to it; and their transcript libraries are exactly the kind of proprietary corpus labs pay for, already sold on subscription.
What would close it. Three or four conversations. Ask a GLG or AlphaSights account director whether any AI lab is a client and what they buy. Ask a lab's data-procurement lead whether they have ever run a request through an expert network. Check whether AlphaSense's Expert Insights product has lab logos.
Difficulty: low, but interview-only. No amount of desk research will produce this. It has never been published because neither side has any reason to publish it.
2. Is Paraform's model exclusive, or five recruiters per role?
Why it matters this much. Paraform's entire proposition to the supply side is that it pays independent recruiters ~70% of a placement fee against the 25–35% they would keep inside an agency. Doubling a recruiter's income is a pitch that needs no product to be believed.
But 70% of a fee you win is a completely different number from 70% of a fee you win one time in five. If a role runs with five recruiters competing, the expected value per role collapses by 80% and the headline split is a marketing number rather than an earnings number. Every claim in the Contingency recruiting marketplaces vertical — supply-side liquidity, leakage resistance, whether the marketplace can hold a recruiter at all — depends on which of these is true.
What we have. Two sources that flatly contradict each other, both of them vendor content. Frontlines.io describes "forced exclusivity"; Dover, a direct competitor, describes five recruiters per role. The Frontlines piece returned 403 on every attempt. Paraform publishes no pricing page and its public terms of service are silent on post-introduction direct hiring. Its actual Recruiting Agreement, which would contain the non-circumvention and exclusivity terms, is not public.
What would close it. Get one recruiter who has worked roles on the platform to describe how a req is assigned. Or get a copy of the Recruiting Agreement. Or open the Frontlines article from a browser rather than a fetcher.
Difficulty: very low. This is a five-minute answer for anyone with one contact in contingency recruiting, and it is the highest value-per-minute question on the page.
3. What does any lab actually pay per expert hour?
Why it matters. Every margin figure in Expert data for frontier labs is derived from one side of the trade. Contributor pay is well documented because recruiting contributors is marketing — Prolific's minimums, Outlier's advertised rates, Surge's disputed $18–30/hr, teleoperators at $25–50/hr. The buy side is documented almost nowhere. The only published take rate in the vertical is Prolific's 42.8%, and Prolific is not selling the same product as Mercor.
Without a buy-side price, the sector's implied margins are computed from a leaked figure for one company (Mercor's 27–33%) and generalised. That is a very thin base for a claim about a $8.5B market.
What we have. Mercor's leaked 27–33%. Prolific's published 42.8%. Handshake's ~60–70% payout implying ~30–40%. Turing's disclosed 50–55% markup implying a 33–36% take. Appen's audited 40.3% and Innodata's ~40% gross margins as the public anchors. Epoch's RL-environment price points, which are the best buy-side data anywhere and cover artefacts rather than hours. Nothing at all for Surge, Scale, Labelbox, Toloka, Snorkel, Pareto or Uber AI Solutions.
What would close it. A single lab-vendor MSA with a rate card in it. Failing that, three contributors on the same project at the same vendor, plus one buyer-side estimate of what the project cost.
Difficulty: high. Rate cards are the most closely held document in this sector, and NDAs run in both directions. Realistically this closes by triangulation, not by disclosure.
4. What is the real take rate at every private company that discloses nothing?
Why it matters. This is question 3 generalised, and it is why the The capital register shows multiples as ranges rather than numbers. Scale, Surge, Turing, Snorkel, Labelbox, Sama, Toloka, Later, Whalar and every expert network report or leak gross. Without net, every multiple in the top half of the register is a five-to-ten-times range rather than a figure.
What we have. Six companies pinned down: Mercor (27→33%, leaked), Prolific (42.8%, published), micro1 (30–40% implied and internally contradictory), Appen (40.3%, audited), Innodata (40%, audited), Handshake (~30–40% implied from payout). That is six out of roughly forty private companies.
What would close it. Nothing wholesale. This closes one company at a time, and mostly through leaks.
Difficulty: structurally hard. The sector has a positive incentive to keep gross as the headline. See GMV is not revenue for why.
5. Did Surge's round close, and at $15B or $25B?
Why it matters. Surge is the second-largest company in the vertical, the only large bootstrapped one, and a $10B ambiguity sits in the middle of the register. At $15B on 2024 revenue the multiple is ~12.5x; at $25B it is ~21x. Those are different judgements about the same company.
More importantly, Surge is the sector's existence proof that this business can be built without venture capital. What price it eventually accepted, and whether it accepted one at all, is the cleanest available read on what the market thinks a profitable, neutral, bootstrapped human-data business is worth.
What we have. Reuters reported a raise of up to $1B at ~$15B (Jul 2025). Bloomberg reported "at least $25B." A LinkedIn post said $30B. No confirmation of a close, size, lead or post-money anywhere. Forbes and Bloomberg subsequently treat Edwin Chen as a multi-billionaire, which implies something priced.
What would close it. A Delaware certificate of incorporation amendment showing the new preferred series and its price. Or a single confirmed press report of the close.
Difficulty: low-to-moderate. This is a filings question, not an interview question, and the filing exists if the round closed.
6. Is the managed clipping agency's spread real?
Why it matters. The pure marketplace take in Paid creators, clipping and UGC ad ops is 9–18% — Whop at 10%, ClipAffiliates at ~18% blended, Mainstage tiering 12% down to 6%. That is a payments fee, not a marketplace margin, and it will not support a venture-scale business.
The claimed money is one layer up: managed agencies that buy views at $0.50–$2 CPM and bill brands at a blended rate plus a retainer. If that spread is real and holds, the managed position is the only defensible one in the vertical — and it is exactly the position the holding companies have been buying (Publicis/Captiv8, Accenture/Whalar). If it is not real, the whole vertical is a 10% payments rail with a fraud problem.
What we have. Nothing. No managed clipping agency discloses its buy-side or sell-side pricing. The highest-margin position in the vertical is entirely undocumented. Circumstantial evidence exists in both directions: Anthony Fujiwara's clipping company reportedly did $7.7M in sales in ten months; Cluely hired 700+ clippers directly, which is the disintermediated case.
What would close it. One agency's rate card plus one campaign's payout ledger. Or the Forbes/Sobrado series, which is the deepest reporting on clipping economics and is 403-blocked.
Difficulty: moderate. The Forbes series is a manual read away. The rate card needs a relationship.
7. What is Distyl's actual ARR?
Why it matters. 58x is the highest multiple in the register, and it is the load-bearing example in Forward-deployed engineering for whether services-as-software works. The denominator — $31M — comes from GetLatka, a source class the atlas explicitly discards after it returned $330K of revenue for a company doing over $1B. The numerator is solid: $175M at $1.8B post, in a press release.
So the most aggressive mark in the sector is being judged against a number from a source the atlas says to throw away. It could be 58x. It could be 20x. Nobody reading this knows.
What would close it. Any second source on Distyl's revenue. A Delaware filing will not give it; a customer count and an average contract value would bound it.
Difficulty: moderate. Private services companies at 111 employees rarely leak revenue, but they do leak headcount and named logos, which bound the range.
8. Is there a measurable leakage rate anywhere?
Why it matters. Disintermediation is the risk every marketplace investor asks about and the one nobody can price. The atlas's formula — leakage ≈ value per relationship ÷ (frequency × switching friction) — is structural reasoning, not measurement, and the page says so.
What we have. One confession and one blocked paper. Upwork's own SEC filing says the number is unmeasurable. Gu & Zhu's Management Science study on freelance-marketplace disintermediation is the one study with a measured rate, and INFORMS blocks automated access to it.
What would close it. Read the paper. It exists, it is measured, and it is one library visit away.
Difficulty: trivial, for anyone with institutional access. This is the highest ratio of importance to effort on the whole page.
9. Does the EU Platform Work Directive actually catch this operator model?
Why it matters. 2 December 2026 is about three months away as this is written. The Directive flips the burden of proof, catches you on where the worker sits rather than where you are incorporated, and closes the BPO workaround with joint-and-several liability. The law is about to arrive argues it is the hardest date on the sector's calendar.
What is not established is how each Member State transposes it, and whether an operator whose workers are genuinely episodic and self-directed falls inside "digital labour platform" as each national law defines it. The Directive sets a floor; twenty-seven implementations sit on top of it.
What would close it. The transposition statutes themselves, as they land. Several will not exist until close to the deadline.
Difficulty: moderate, and time-dependent. This closes itself over the next six months whether anyone works on it or not.
10. What happens to Mercor's margin as it scales?
Why it matters. The leaked figure is 27% rising to 33%. Rising margin during hypergrowth is the single strongest piece of evidence for the services-as-software thesis inside the atlas's own data — it is the migration Sequoia's argument requires, actually observed. It is also six percentage points of one leaked number about one company over an unstated period, and it could as easily be mix shift as operating leverage.
What would close it. A third and fourth data point on the same series.
Difficulty: high. Depends on another leak.
Part 2 — The thirty-three unresolved contradictions
These are places where two sources the atlas relies on disagree, and neither was resolved. They are stated here in the reader's terms rather than the researcher's. Where a contradiction is load-bearing on a page, that page carries a :::gap callout pointing here.
About revenue and margin
1. Mercor's $614M is gross, and net on the same base is also about $600–660M. Two entirely different numbers that happen to look identical: $614M gross for H1 2026, and net on a $2B gross run-rate at a 33% margin. They must never be collapsed, and the coincidence makes collapsing them easy.
2. micro1's reported economics are arithmetically impossible. The same coverage says it retains 60–70% of gross and that net is $150–200M on $500M gross — which is a 30–40% take. Both cannot hold. The atlas uses the net figure and treats the retention rate as weak.
3. Surge pays contractors either $18–24/hr or about $30/hr. Sacra's 30–40 cents per working minute against a widely reported "around $30/hour." A ~50% discrepancy on the number that drives Surge's implied margin.
4. Scale was profitable in H2 2025, or was not. Scale's own blog says the data business was profitable in the second half of 2025; Business Insider reported in July 2025 that it was not. Different periods, never stated as such by either.
5. Handshake's legacy business is ~$150M or $190M, and its AI team grew from 15 to 150 or from 3 to 150. Small in isolation; it matters because the group net-revenue figure is built by subtraction.
6. Appen's currency basis is blended in the underlying research. US$230.8M and US$232.7M for FY2025, a TTM figure quoted in A$361.9M, and a peak given as both US$4.3B and ~A$4B+. The USD and AUD series must never be mixed, and in the raw notes they were.
7. Turing's "roughly tripled" describes $120M → $300M, which is 2.5x.
8. The $8.5B market-size cross-check adds gross figures to net figures. The bottom-up sum that corroborates the sector total uses Mercor's gross beside Appen's net. The total is roughly right by luck, not by construction.
9. ShopMy's take rate does not reconcile with its revenue. Published rates are 2.9% direct and ~3.9% subaffiliate; $80M of net on >$1B of GMV implies a blended ~8%. One of those is wrong, or there is a revenue line nobody has described.
10. Whop's take rate does not reconcile either. $142M net annualised at a ~5.5% blended take implies ~$2.6B of annualised GMV — but cumulative lifetime GMV was $2.67B by February 2026. Annualised cannot be within 3% of lifetime unless one of the two figures means something other than what it says.
11. The "92x ShopMy" figure does not exist. It appears in the raw research and is arithmetically impossible. The real ratios: GMV to net ≈ 12.5x, valuation to net 18.8x, valuation to GMV under 1.5x. A multiple on GMV is always lower than a multiple on net. The error is still live in the source file that generates it.
12. ColdIQ's "$7M ARR" is gross agency billings, and it sits in a take-rate league table beside net fees. Not comparable, and the atlas has seen it presented as though it were.
13. Superside's implied take rate is arithmetic on Glassdoor pay data, and the two bands derived from the same two inputs — 70–88% and 55–65% — differ by about thirty points. That is not a range, it is a disagreement with itself.
About valuations, prices and dates
14. Surge's 2026 round: Reuters ~$15B, Bloomberg "at least $25B", never confirmed closed. See question 5 above.
15. Meta paid $14.3B or $14.8B for 49% of Scale. Both figures circulate and are used interchangeably across the research. The atlas standardises on $14.3B and flags it.
16. Distyl's 58x rests on a GetLatka ARR estimate — from the source class the research explicitly says to discard. See question 7.
17. Toptal's ~$167M revenue comes from the same discredited estimator class. Toptal is described in the atlas as its single strongest data point, and the revenue attached to it is unverifiable.
18. Publicis paid $175M for Captiv8, or an undisclosed sum. ContentGrip gives a price; every other source says terms were not disclosed.
19. Accenture's Whalar deal is dated 10 June 2026, or announced 8 June and closed 31 July 2026. The atlas uses the second.
20. Andela raised ~$330M or ~$381M, and is either a "zombie since 2022" or a company with a documented September 2024 CEO appointment, January 2025 layoffs and a January 2026 Woven acquisition. The second set of facts is better sourced; the first framing is more widely repeated.
21. Scale's current CEO is unreconciled: Jason Droege from June 2025, Francis deSouza named July 2026, with a claim of three CEOs in fourteen months. The count and the sequence do not quite fit together.
22. The Meta CPM baseline is used two ways. $14.68 is the conversion-objective CPM; $14.19 is a general Facebook CPM. They are used interchangeably in the clipping cost comparison, which is exactly the comparison that argument turns on.
About concentration and sourcing
23. Appen's five-client concentration is two different facts presented as one. 80% at the 2020 peak, and 74.3% (up from 67.3%) in Appen's own FY2025 report. Different years, different companies in the numerator, routinely conflated into a single trend.
24. The GLG S-1 is both read and unread. One research file cites it as fetched, with specifics — top-ten clients at 19.0%, 90% recurring revenue, 22 of the top 25 clients retained. Another calls the same document "the single highest-value unread document." Both cannot be true, and the specifics look like a genuine read.
25. Palantir's gross margin is internally inconsistent. 82% (84% ex-SBC) from the FY2025 10-K in one section; 84.8% TTM in the market-comparison table in the same file. Same company, same file, two numbers.
26. The Foody/Nvidia concentration comparison has no retrievable citation. One sentence in the research asserts that Brendan Foody publicly compared Mercor's customer concentration to Nvidia's (four customers = 61% of revenue), with no source attached. The comparison is repeated in the atlas because it is illuminating, which is precisely the wrong reason to repeat something.
About the legal and operational surface
27. The working-capital arithmetic funds only the COGS portion across the payment gap, while describing the result as the whole receivable being tied up. The two differ by about 5.5 points of revenue at net-60, which is material to anyone sizing a facility. You pay weekly, they pay in sixty days separates them explicitly; the underlying research did not.
28. The FTC civil-penalty figure differs between files: $53,088 citing 90 FR 5580, against a $51,744–$53,088 range elsewhere. The 2026 inflation adjustment could not be confirmed either way.
29. No factoring or receivables-financing rate exists anywhere in the research. The ~14.9% annualised cost of 2/10-net-60 on the working-capital page is derived arithmetic, not a sourced rate. Nobody in this sector has published what it actually costs to finance a labour receivable.
30. Data-centre labour has demand evidence and no vendor evidence at all. Every figure on that page describes the buyer or generic construction staffing. The $14.2M-per-month delay cost is single-source and single-project-size, and no intermediary in that vertical has published anything.
31. Crunchbase's Q2 2026 global early-stage figure is garbled — "$589 billion" against a $205B quarter — and the AI share of venture runs 60% (Carta) / 80% (Crunchbase) / 86% (PitchBook). The atlas quotes the range rather than picking one, and does not use the $589B figure at all.
32. Paraform's exclusivity model. See question 2. It is both a contradiction and the second-most-important open question on this page.
33. The take-rate league table in the raw research mixes bases. Expert networks at 70–80% (a genuine take on a call fee), Clay agencies at 55–75% (a gross margin), compute brokerage at "low single digits" (a fee), and red-teaming at ">90% effective" (an inference about crowd output monetised as software). These are four different quantities in one column. What a rake can actually be rebuilds the comparison on a consistent basis; the raw table should not be quoted.
Part 3 — Structural gaps: where the evidence does not exist at all
The contradictions above are disagreements between sources. These are holes where no source exists — and that distinction matters, because these will not close with better search.
No quantitative leakage rate exists anywhere
The single most important missing number in the sector. Every claim about disintermediation — in this atlas and in every pitch deck it describes — is structural reasoning plus one confession. Upwork, after twenty years and with complete transaction data, states in its filings that the number is unmeasurable. The one academic study with a measured rate sits behind a publisher that blocks automated access.
The consequence: nobody can price the risk that the two sides route around you. Investors ask about it, founders answer with adjectives, and no one on either side of the conversation has a number. Getting cut out is the atlas's lowest-confidence structural page for this reason.
Latin American platform-work law
No usable 2025–26 source was found for Brazil, Chile, Colombia or Mexico. That is a real hole, because LatAm is one of the two or three geographies an operator would actually pick for cost-and-timezone reasons, and it is where the atlas has the least to say about legal exposure. Where the supply can legally live covers the EU, UK, India, Kenya and the Philippines and goes quiet on a continent.
No documented crowd-worker data leak
Every operator in this sector runs NDAs at scale over a distributed crowd touching frontier-lab data, and every lab imposes a security bar in its MSAs. No published incident of a crowd worker leaking lab data was found. The risk is universally assumed and nowhere evidenced.
Two readings, and the atlas cannot distinguish them: either the controls work, or the incidents are absorbed commercially and never surface. Both are consistent with the record. The absence is not reassurance.
No factoring rate for labour receivables
See contradiction 29. A labour middleman is structurally a bank lending its customers money at 0%, and there is no published price for the debt that funds it. The ~14.9% figure on You pay weekly, they pay in sixty days is the annualised cost of an early-payment discount, computed by the atlas, standing in for a rate nobody publishes.
The $1M–$20M bootstrapped band
The register jumps from Mediamaxxing at "no funding found" to Distyl at $202M raised. The band between — genuinely bootstrapped operators doing $1M to $20M — is almost entirely missing, and its absence is a retrieval artefact, not a finding. Indie Hackers, X threads, small-agency podcasts and founder-community posts are not reachable through the surfaces this research used, and the estimators that claim to cover this band failed their spot checks.
This is the band most relevant to anyone actually starting one of these businesses, and the atlas has the least to say about it. It is also the band where the Surge and Toptal pattern — refuse capital, grow slower, keep the company — would show up if it generalises.
Buy-side pricing in voice and speech data
Voice, speech and low-resource language data has the best supply-side data in the atlas and the worst buy-side data, for a structural reason: contributor rates are advertised because recruiting contributors is marketing, while every buyer is routed through "get in touch." You can find out to the cent what a person is paid to read a paragraph in Yoruba. No vendor in that vertical publishes revenue or a rate card. The 60–85% spread on that page is a working assumption, and the page says so.
Lab procurement: who signs, and at what threshold
No source names a role. The reporting implies research-led budget authority at the labs rather than procurement-led, but nobody states a title or an approval threshold. For a vendor building a sales motion, this is the most operationally useful missing fact in How much money is actually in the buyer pool, and it is interview-only.
Non-compute operating spend at any lab
OpenAI's inference costs are public. R&D headcount, stock compensation and vendor spend are not, at any lab. Every human-data budget figure in the atlas is therefore a leak, an estimate or an aggregate — never a disclosure.
Vendor-side evidence for the VC channel
No source quantifies what share of any vendor's revenue actually originates from a VC or accelerator referral. Selling through the investor argues from the buy side — HubSpot's published discounts, AWS credits, Vanta's YC penetration — because the sell side is undocumented. "Our investor will introduce us to their portfolio" remains unproven revenue in both directions.
Non-payment and bad-debt rates when selling to startups
62% of seed-funded startups eventually fail, and no dataset attributes vendor churn to customer death, nor gives a bad-debt rate for selling into that population. How much money is actually in the buyer pool can tell you the customer is mortal and cannot tell you what that costs.
Part 4 — What would be worth doing next, in order
Ordered by value divided by effort, not by importance alone.
1. Read the Gu & Zhu leakage paper. One library visit. It closes the atlas's largest theoretical hole with an actual measured number, and it is the only item on this list that converts a structural gap into a fact in an afternoon.
2. Open the six blocked articles by hand. Frontlines.io on Paraform's exclusivity; the Forbes/Sobrado clipping series; The Information on Handshake's revenue and on eval-startup revenue; Reuters and Bloomberg on Surge's round. A browser and an hour. This alone would resolve four of the thirty-three contradictions and two of the top six questions.
3. Ask three people the expert-network question. One network account director, one lab data lead, one AlphaSense customer. If the answer is "no," the most valuable arbitrage in the atlas is confirmed open.
4. Re-source Distyl's ARR. The highest multiple in the register currently rests on a discredited estimator. Any second source — a headcount-derived bound, a named-logo count, a customer's spend — collapses the uncertainty on the atlas's flagship services-as-software example.
5. Pull the GLG S-1 properly and settle contradiction 24. It is on EDGAR, it is free, and it contains the only audited P&L for the most durable version of this model. It would also give a real 2018–2021 revenue series against which the "$3B market" claim can be checked.
6. Check for a Surge Series A certificate of amendment in Delaware. Closes the $10B ambiguity in the middle of the register, or confirms that no round closed — which is itself the more interesting answer.
7. Get one lab-vendor MSA with a rate card. Hardest item that would change the most. It converts the entire margin discussion in Expert data for frontier labs from inference to measurement.
8. Build a historical multiples panel — 2015, 2019, 2021, 2026. Every multiple in What the public market pays for labour is a single 2026 snapshot. Robert Half at a 0.2% operating margin is clearly at a cyclical trough; Fiverr at 0.09x EV/revenue may be at a sentiment trough. A four-point panel separates structural from cyclical, and it is pure desk work on data that already exists.
9. Read the complaints behind the Mercor, Schuster and Belardi dockets. The atlas knows what was filed and not what is alleged. The theories matter a great deal for anyone designing a contributor agreement.
10. Find the $1M–$20M band. Different retrieval surface entirely — founder communities, niche podcasts, direct outreach. Slowest item here and the one most likely to change what an operator actually does on Monday.
Everything above is a list of things the atlas does not know. It is deliberately long, and it should be read as a caution about the pages that do not carry these warnings as much as about the ones that do. A page with confidence: high on it is high-confidence relative to the rest of this record, not relative to diligence. See How this was built.