A bibliography tells you what was read. This tells you what is load-bearing, and what happens to the argument above it if the claim turns out to be wrong.
The hierarchy used
Sources were ranked in this order, and the ranking is doing real work — several of the atlas's most quoted numbers sit near the bottom of it.
- SEC filings and audited annual reports. A 10-K carries auditor liability and a revenue-recognition policy you can read. Every take-rate benchmark in What a rake can actually be and every public multiple in What the public market pays for labour comes from this tier, which is why those two pages carry
confidence: highand most vertical pages do not. - Regulations and primary legal text. EUR-Lex, the Federal Register, IRS guidance pages. Unambiguous, dated, and they say what they say.
- Court dockets. Verified through CourtListener's RECAP API: real cases, real filing dates, real nature-of-suit codes. What they do not give is the complaint, so the atlas can say a fraud claim was filed and cannot say what it alleges.
- Reputable press with a named reporter. TechCrunch, NYT, Digiday, NPR, Guardian, Fortune, Reuters, Bloomberg. Downgraded one notch where the article itself could not be opened — see the closing section.
- Leaks reported by press. Mercor's 27–33% gross margin, Anthropic's RL-environment budget. These are the most valuable numbers in the atlas and the least verifiable ones. Always attributed as leaks in the sentence.
- Analyst profiles (Sacra, Contrary, Dealroom). Useful, frequently the only source for a private company, occasionally wrong. Treated as a single secondary source, never as a disclosure.
- Company blogs and pricing pages. Excellent for prices, which a company has an interest in stating accurately, and for funding announcements. Worthless for market size, competitor comparison or anything with the word "leading" in it.
- Vendor content marketing. A blog published by someone selling the thing it measures.
- Inference. Arithmetic the atlas did itself, marked as derived wherever it appears.
Because a circular citation is worse than no citation. Several 2026 "benchmarks" in the creator and robotics verticals trace back through two or three intermediaries to a vendor blog that cites a second vendor blog that cites the first. The number acquires the appearance of independent corroboration purely by being repeated, and there is no underlying measurement anywhere in the chain.
At least the atlas's own arithmetic shows its inputs. A vendor benchmark that cannot be traced to a measurement is not a weaker fact than an inference — it is a different kind of object, and the honest move is to name the vendor in the sentence so the reader can discount it themselves. That is why Paid creators, clipping and UGC ad ops, Design, video and content production, Voice, speech and low-resource language data and Adversarial evals and red-team crowds all carry low confidence despite having plenty of numbers on them. Numbers are not evidence.
The register
Source type uses the tiers above. Confidence is in the evidence, not in the conclusion drawn from it.
| Claim | What it supports | Source type | Confidence | Where used |
|---|---|---|---|---|
| Robert Half contract staffing ran a 39.0% gross margin in FY2025; perm placement ran 99.8% | The entire COGS test — selling hours versus selling a matched outcome, inside one company | SEC filing (RHI FY2025 10-K) | High | What the public market pays for labour, What a rake can actually be, Marketplace, staffing firm, BPO or agency |
| Upwork FY2025: GSV $4.028B, marketplace revenue $682.9M, take rate 18.7% | The best public marketplace take-rate benchmark and the GSV-to-revenue ratio | SEC filing | High | What a rake can actually be, GMV is not revenue |
| Upwork books its marketplace net as agent and Managed Services gross as principal, and says so | The ASC 606 principal/agent test, demonstrated inside one company | SEC filing | High | GMV is not revenue, Marketplace, staffing firm, BPO or agency |
| Upwork's risk factors flag losses from clients failing to pay "particularly where we advance payments to talent" | That the working-capital gap is a real, disclosed risk and not a theoretical one | SEC filing | High | You pay weekly, they pay in sixty days |
| Public EV/revenue: Accenture 1.56x, Robert Half 0.84x, ManpowerGroup 0.22x, Fiverr 0.09x, Palantir 71.2x | The ~1.6x ceiling on labour businesses and the three documented escapes from it | Market data (stockanalysis.com, 27–29 Aug 2026) | High | What the public market pays for labour |
| Palantir FY2025 gross margin 82% (84% ex-SBC) | That FDE-heavy delivery can still carry a software margin if the artefact sold is a licence | SEC filing (PLTR 10-K) | High — but the atlas's own notes give 82% and 84.8% in two places | Forward-deployed engineering, What the public market pays for labour |
| C.H. Robinson: ~8.5% gross margin on $17.0B of gross freight revenue | That a gross-booking business restates to a software-looking multiple with no economic change | SEC filing + market data | High | GMV is not revenue |
| Uber Mobility take 30.4%, Delivery 19.0%; Etsy 24.2%; DoorDash 13.4%; Airbnb 13.4% | The empirical ceiling on a pure discovery rake at scale | SEC filings (FY2025 10-Ks) | High | What a rake can actually be |
| GLG S-1 (Oct 2021): $589.1M FY2020 revenue, >70% segment contribution margin, top-ten clients 19.0%, no client above 6% | That a 70%+ take can be structurally durable, and that the survivor of this model deliberately fragmented its buyers | SEC filing | High — but the atlas's own notes both cite it as fetched and call it unread | Expert networks, GLG, One customer is a binary event |
| Appen FY2025: US$230.8M operating revenue, 40.3% gross margin, US$12.2M underlying EBITDA, top-5 clients 74.3% (from 67.3%) | The only audited P&L for a pure-play in this sector | Audited annual report | High | Appen, One customer is a binary event, The capital register |
| Appen revenue US$447.3M (FY2021) → US$232.7M (FY2025); market cap A$330M; 97% drawdown from a ~US$4.3B peak | The base rate for what a concentrated data vendor is worth after one hyperscaler leaves | Market data + press | High | Appen, One customer is a binary event, The capital register |
| Alphabet terminated a ~US$83M Appen contract on 22 Jan 2024; shares fell 40–41% in a day | That customer concentration is a cliff, not a slope | Reputable press | High | Appen, One customer is a binary event |
| Innodata TTM revenue $317.2M (+39%) at ~40% gross margin, $1.95B market cap = 6.2x | The upper bound the public market will pay for an AI-data business it can see inside | Market data + company release | High | What the public market pays for labour, The capital register |
| Directive (EU) 2024/2831 must be transposed by 2 December 2026; it flips the burden of proof and applies by where the worker sits | The hardest date on the sector's calendar and the whole EU exposure argument | Regulation (EUR-Lex, primary text) | High | The law is about to arrive, Where the supply can legally live |
| Prolific publishes a 42.8% platform fee for corporate customers (33.3% academic), on top of participant rewards | The only published, unambiguous take rate in the expert-data vertical | Company pricing page | High | Prolific, What a rake can actually be, Expert data for frontier labs |
| Whop takes 10% of Content Rewards payouts | The market-clearing price for running clipping rails, and the cap on the layer below managed agencies | Reputable press (Digiday), corroborated independently | High | Whop, Paid creators, clipping and UGC ad ops, What a rake can actually be |
| Clipping pays $1–$5 per 1,000 views; AI startups pay ~$25, MLB $1, Polymarket $0.50 | That "clipping is cheap" is false for exactly the buyer with venture money | Reputable press (Digiday, NPR) | High | Paid creators, clipping and UGC ad ops |
| Global venture H1 2026 $510B; OpenAI + Anthropic took $217B = 43% of it | That the headline TAM is not addressable and the buyer pool splits in two | Reputable press (Crunchbase) | High | How much money is actually in the buyer pool |
| HubSpot for Startups: 4,000+ approved partners, 35,000+ startups, 90%/50%/25% discount | That the VC channel is a discount programme with a price, not a recommendation engine | Company page | High | Selling through the investor |
| Deel: contractor management from $49/contractor/month, EOR from $599/employee/month | That per-head compliance tooling is priced for staffing and unusable at crowd scale | Company pricing page | High | Where the supply can legally live, The law is about to arrive |
| FTC civil penalty per violation $53,088 (90 FR 5580) | The scale of endorsement-disclosure exposure | Regulation (Federal Register) | High — but a $51,744–$53,088 range appears elsewhere in the same research | Paid creators, clipping and UGC ad ops |
| RL-environment pricing: contracts six-to-seven figures per quarter; website replica ~$20k, Slack clone ~$300k; tasks $200–$2,000; exclusive deals price at 4–5x non-exclusive | That labs are increasingly paying for denial, not data — the most commercially interesting number in the vertical | Research-organisation analysis (Epoch AI) | High | Expert data for frontier labs, What better models do to each layer, Mechanize |
| Meta paid $14.3B for 49% non-voting of Scale at a $29B valuation, plus Wang and a research team | The template for licence-and-hire, and the natural experiment in why neutrality is the product | Reputable press (NYT) | High — though the price circulates as both $14.3B and $14.8B | Scale AI, One customer is a binary event |
| Sama cut 1,108 Nairobi jobs (Apr 2026) after Meta ended its contract | Appen's failure mode reproduced one layer down the supply chain | Reputable press (Guardian, Washington Post, TechCabal) | High | The capital register, Where the supply can legally live |
| iMerit sold to EXL for up to $310M (announced Jun 2026, closed Aug 2026) | The best disclosed exit price for a real operating data business | Reputable press | High | The capital register |
| Adecco bought Vettery for $100M (2018); Hired raised $132.7M and was absorbed for an undisclosed sum | What a failed hiring marketplace is actually worth to a strategic | Reputable press | High | Triplebyte and Hired, The capital register |
| Court dockets: Ramey v. Scale (FLSA), Schuster v. Scale (personal injury), Scale v. Mercor (Defend Trade Secrets Act), Belardi v. Appen, and a Mercor wave of fraud/contract suits Apr–May 2026 | That the modal legal claim in this sector is now four different things, not just wage-and-hour | Court docket (CourtListener RECAP) | Medium — cases and codes verified; complaints not read | The law is about to arrive, Who is actually on the other end |
| Mercor's gross margin is 27% rising to 33% | The single most important margin number in the atlas — the only real gross margin any expert-data marketplace has produced | Leak reported by press | Medium | Mercor, Expert data for frontier labs, GMV is not revenue |
| Mercor contractors take 60–70% of gross billings | The 13x-vs-38x restatement, and by extension the whole gross-vs-net argument | Analyst profile (Sacra) | Medium | GMV is not revenue, The capital register |
| Mercor at $2.0B gross annualised (Jun 2026), from $1M in early 2023 | That the budget is real and the growth is not disputed | Reputable press + analyst profile | High | Mercor, Expert data for frontier labs |
| ~91% of Mercor's revenue comes from foundation-model companies | The concentration case against the whole cohort | Press | Medium | One customer is a binary event, Mercor |
| Anthropic discussed spending over $1B on RL environments in a year | That the budget the atlas describes exists at the scale claimed | Leak reported by press (The Information via TechCrunch), corroborated by Epoch | Medium | Expert data for frontier labs, How much money is actually in the buyer pool |
| Google spent ~$150M with Scale in 2024, planned ~$200M in 2025 | The only vendor-specific lab spend figure anywhere | Analyst profile | Medium | How much money is actually in the buyer pool, Scale AI |
| Vertical aggregate ~$8.5B/yr gross across 50+ vendors, top four taking 75%+ | The market-size figure the atlas uses, with a bottom-up cross-check | Analyst synthesis of a vendor map | Low — the cross-check adds gross figures to net figures | Expert data for frontier labs |
| Handshake AI at ~$1.0B gross annualised (Apr 2026), ~$300M net; group $1.10B gross / ~$450M net | That the fastest zero-to-$1B in the vertical happened on pre-existing supply, and looks cheap only because the mark is stale | Press (headline-level) + analyst profile | Medium | Handshake AI, The capital register |
| micro1 at a $500M gross run-rate with net at $150–200M | The small-cap growth case, and a worked gross-vs-net restatement | Reputable press (TechCrunch, Dataconomy) | Low — the same source also says it retains 60–70%, which is arithmetically incompatible | micro1, What we could not establish |
| Turing's model is "a staffing spread"; $300M+ annualised 2024, profitable — effectively gross | That a $2.2B mark sits on a gross staffing number | Analyst profile + company PR | Medium — nothing found for 2025 or 2026 | Turing, GMV is not revenue |
| Invisible: $134M revenue 2024 (+123%), ~$15M EBITDA (11% margin) | The only disclosed profit figure in the private cohort | Analyst profile | Medium | Invisible Technologies |
| Surge reached >$1B revenue in 2024 — basis unstated, almost certainly gross — bootstrapped, ~130 FTEs, Chen owns ~75% | The strongest counterexample to the capital thesis in the atlas | Reputable press (Inc., The Verge) + analyst profile | Medium | Surge AI, The capital register |
| Surge's 2026 round: Reuters ~$15B vs Bloomberg "at least $25B", no confirmed close | Nothing — it is the largest single ambiguity in the register | Press, conflicting, headline-only | Low | Surge AI, What we could not establish |
| Surge contractor pay: 30–40c per working minute ($18–24/hr) vs "around $30/hour" | Surge's implied margin, which cannot be computed while this is open | Analyst profile vs press, ~50% apart | Low | Surge AI, What we could not establish |
| Scale cut 200 FTEs and ~500 contractors in Jul 2025; rival labs withdrew within weeks of the Meta deal | The neutrality argument — the closest thing to a controlled experiment the sector has | Press (Computerworld, TechCrunch) | Medium — the outcome is reported everywhere, the causal chain is sourced nowhere | Scale AI, One customer is a binary event |
| Expert networks: an expert at $200/hr bills the client $800–$900 (range $700–$1,100) | The 70–80% take that anchors the whole "a durable rake is possible" argument | Industry forum post, corroborated by an independent pricing blog | Medium | Expert networks, What a rake can actually be |
| Expert-network market ~$3B (2025), growing ~12%/yr; GLG's share fell 51% → 24% over a decade | That the supply is not the moat even in the most durable version of this model | Industry blog (single source) | Low–Medium | Expert networks, GLG |
| No evidence found that any expert network sells expert time or transcripts to frontier labs | An absence, and the single most valuable unanswered question in the atlas | Absence of evidence after directed search | Low (as an absence) | Expert networks, What we could not establish |
| Paraform pays independent recruiters ~70% of the placement fee against 25–35% inside an agency | The supply-side wedge the entire recruiting vertical is built on | Vendor content marketing — a competitor's blog and a third-party comparison site; Paraform publishes no pricing page | Low | Paraform, Contingency recruiting marketplaces |
| Paraform has raised $65M and paid $50M to recruiters | The derived ~$71M lifetime gross / ~$21M lifetime net that sizes the business | Company blog (both figures, same announcement) | Medium for the inputs; the arithmetic is inference | Contingency recruiting marketplaces, Paraform |
| Paraform enforces "forced exclusivity" — versus five recruiters per role | Whether the 70% split is real earnings or headline earnings. Nothing decides more per word | Vendor content marketing on both sides, flatly contradictory; the primary source 403'd | Low | What we could not establish, Paraform |
| US staffing market $180.2B (2026); temp/contract is 89%, perm placement 11% | That the recruiting TAM is off by an order of magnitude for a perm-placement marketplace | Industry press citing SIA/ASA | Medium | Contingency recruiting marketplaces |
| Contingency fees 15–25% of first-year salary; retained search ~33% | The fee ladder the marketplace take is carved out of | Industry convention; no primary source obtained | Low | Contingency recruiting marketplaces, Sales talent: AEs, SDRs and the people who close, What a rake can actually be |
| Vanta: $300M ARR (Apr 2026), 16,000 customers, ~$19K ACV, three-quarters of YC | The empirical price point for a horizontal service sold to funded startups at scale | Analyst profile | Medium | How much money is actually in the buyer pool, Compliance, finance and back office |
| SOC 2: audit $10k–$50k once; Vanta/Drata $20k–$80k a year | That the margin has already migrated from the licensed human to the software above them | Vendor content marketing (a competing compliance vendor) | Medium — the prices are checkable, the framing is not neutral | Compliance, finance and back office |
| Gray Swan's UK AISI agent challenge had a total prize pool of $171,800 | That four of the best-funded organisations on earth bought a coordinated attack for less than one senior researcher's loaded cost | Archived challenge brief | Medium | Adversarial evals and red-team crowds, Gray Swan AI |
| Gray Swan's revenue comes from Shade and Cygnal, not from the Arena crowd | The >90% effective take, and the only clean example of monetising crowd output as software | Reputable press (Forbes Australia) | Medium — no revenue figure exists to check it against | Adversarial evals and red-team crowds, Gray Swan AI |
| Teleoperation bills the buyer $50–$200/hr against operator pay of $25–$50/hr | The 40–65% spread that makes robotics data the atlas's best-scored labour vertical | Vendor content marketing, two roughly converging sources | Low–Medium | Robotics teleoperation and physical-world data |
| Build AI released 1M hours of egocentric factory footage free in Apr 2026, collapsing a market that traded at "a few dollars per hour" | The clearest documented price collapse in the atlas | Vendor content marketing (a competing vendor's blog) | Low | Robotics teleoperation and physical-world data, What better models do to each layer |
| Superside's implied take rate is 70–88%, or 55–65% on generous assumptions | The widest apparent spread in the atlas, attached to the weakest defence | Inference from a vendor pricing page and Glassdoor self-reported pay; the notes' own two bands differ by ~30 points | Low | Design, video and content production |
| A delayed 60MW data-centre project costs its owner $14.2M a month; 349,000–499,000 additional workers needed in 2026 | The urgency of the data-centre buyer, and the entire case for that vertical | Single vendor report, single project size | Low | Data-centre labour and site brokerage |
| Invalid traffic was 18.12% of impressions in Q1 2026 | The bot-contamination discount on any CPM-priced creator campaign | Industry measurement, general programmatic — not clipping-specific | Low as applied | Who is actually on the other end, Paid creators, clipping and UGC ad ops |
| Meta conversion CPM $14.68; TikTok Spark Ads $3.54 vs $4.73 platform average | The comparison the clipping pitch rests on | Vendor content marketing (one vendor sells UGC) | Low | Paid creators, clipping and UGC ad ops |
| Christina Chapman sentenced to 8.5 years (Jul 2025) for a DPRK laptop farm; further US and Ukrainian sentences 2026; OFAC designations Mar 2026 | That state-sponsored infiltration is a hiring-controls line item, not an exotic risk | Press, multiple outlets, dates verified | Medium | Who is actually on the other end |
| ASC 606 control test: primary responsibility, inventory risk, price discretion | Whether an operator books gross or net, which sets the multiple | Accounting standard via Upwork's applied disclosure — the codification text itself was not opened | Medium | GMV is not revenue, Marketplace, staffing firm, BPO or agency |
| 1099-NEC threshold raised from $600 to $2,000 for 2026 payments; several states did not follow | Payouts compliance design | Law-firm and trade commentary; no IRS primary source obtained | Low | The law is about to arrive, Where the supply can legally live |
| A non-US person performing services entirely outside the US earns foreign-source income → no 30% withholding, no 1042-S | The whole cross-border payout design for a European operator | IRS sourcing rule (primary) + inference for the 1042-S conclusion | Medium for the rule, Low for the conclusion | Where the supply can legally live |
| Single-player mode solved cold start for 34% of the 100 largest marketplaces, at ~10:1 revenue-to-funding | The capital-efficiency case for building one side first | Independent blog analysis of 100 companies | Medium | Which side you build first |
| Sequoia, "Services: The New Software" (Julien Bek, 5 Mar 2026): six services dollars per software dollar; intelligence work vs judgement work; autopilots capture the labour budget | The explicit bull thesis the register's multiples are underwritten against | Investor essay — a thesis, not a fact | High as to what it says | The capital register, What the public market pays for labour, What better models do to each layer |
| No measured leakage rate exists anywhere. Upwork's own filing says it is unmeasurable; the one academic study is behind a publisher that blocks automated access | That every claim about disintermediation in this sector, including the atlas's, is structural reasoning rather than measurement | Absence + SEC filing (the confession) | High as to the absence | Getting cut out, What we could not establish |
What is wrong with the sources themselves
Three source-class problems recur, and they are not incidental — they shape which facts the atlas has.
Third-party revenue estimators fail spot-checks badly enough to be unusable
The atlas discards GetLatka-class estimators entirely, on the basis of a spot check it failed catastrophically: its Surge AI page returns $330K of revenue and 3 employees for a company that did over $1B in 2024 with about 130 employees. That is not a rounding error, it is a wrong company. The spot-check is carried as recorded in the research notes; the URL of that particular Surge AI page is not in the record, so this claim goes uncited by design — the alternative, pointing at a different GetLatka page as if it showed the Surge figures, is the kind of near-miss citation this page exists to catch.
The problem is that these estimates propagate. Three figures in the register — Distyl's $31M ARR, Toptal's ~$167M, Superside's $44.9M — come from this source class (GetLatka) and nowhere else. Distyl's is the worst case, because it is the denominator of the highest multiple in the atlas: the "58x" headline rests on an estimate from a source the atlas otherwise says to throw away. Every one of these is tagged [WEAK] on the page where it appears, but a tag is not a fix. If any single number in this atlas should be re-sourced first, it is Distyl's ARR.
Paywalls and bot-blocking mean several facts are sourced from a headline and a dateline
Forbes, The Information, Bloomberg, Reuters, WSJ and Business Insider either 403 automated fetching or sit behind a paywall. Where the atlas needed them, the fact was captured from a headline, publisher and date via news aggregation — and that is all. The claim exists; the reporting behind it was not read.
The concrete casualties:
- Boaz Sobrado's Forbes series on clipping economics (four or more pieces, Feb–Jul 2026) is the deepest reporting on that vertical anywhere and is entirely inaccessible. Paid creators, clipping and UGC ad ops is the atlas's lowest-confidence vertical largely because of it.
- The Information's "Revenue Lags at AI Evaluation Startups" (Apr 2025) is the one trade-press piece squarely questioning the eval-sector marks. It is cited by headline only.
- The Information's Handshake revenue reporting (Apr 2026) is the source for the $1.10B gross figure. Headline only.
- Reuters at ~$15B and Bloomberg at "at least $25B" on Surge's round — a $10B disagreement, both inaccessible, neither confirmable.
- Frontlines.io on Paraform's "forced exclusivity" — the primary source for the single most consequential mechanical question in Contingency recruiting marketplaces — returned 403 on every attempt.
Google News RSS URLs stand in for articles
Where a publisher blocked direct fetching, several citations in the underlying research resolve to a Google News search URL rather than to the article. That is an honest citation of what was actually retrieved — headline, publisher, date — and it is not a citation of an article anyone read.
Roughly a fifth of the register's rows have at least one such citation behind them. They are marked in the research notes and the affected facts are stated with attribution in the sentence wherever they appear on a page. Anyone acting on one of them should open the underlying article first.
The compound effect of these three problems is systematic, not random. It biases the atlas toward SEC registrants, regulations, court dockets, company pricing pages and sites that permit automated fetching, and away from paywalled investigative journalism about private companies. That is precisely the reverse of where the interesting facts about private companies live. See How this was built for what that means for how much weight any of this will take.