Capability Gap

All verticals

Adversarial evals and red-team crowds

Labs pay five and six figures per model to be attacked; the attackers are paid in prize money. The widest spread and the scarcest supply in the sweep.

buildmedium confidence8 minupdated 2026-08-29evals · red-teaming · ai labs · security · crowds
Who is buying
Frontier labs buying pre-release safety evidence; enterprises buying agent security against the EU AI Act and NIST AI RMF
Who is selling
Elite jailbreakers and bug-bounty researchers — a small, self-selecting, competitive population
The spread
Effective take on crowd output >90% [UNVERIFIED — inferred, no revenue disclosure]
Size of the pool
No verified figure. One vendor blog claims $2.26B for 2026 [WEAK]
Read
A brand-new budget line with no salary to be benchmarked against, supply that cannot be recruited by job ad, and a middleman that monetises the crowd's output as software rather than reselling its hours.
Budget depth
5
Supply difficulty
5
Spread
4
Holdability
3
AI direction
5
Speed to first dollar
4

The UK AI Security Institute ran an agent red-teaming challenge on Gray Swan's Arena with OpenAI, Anthropic and Google DeepMind participating. The total prize pool was $171,800 — $86,000 for the leaderboard, $45,000 by volume, $25,800 in first-break bounties at $300 each, $10,000 for over-refusal, $5,000 for new users, minimum payout $100 (MBG Security's archive of the challenge brief). Four of the best-funded organisations on earth bought a coordinated assault on their frontier models for less than the loaded cost of one senior researcher.

That crowd's work sits inside a company that raised $40M at roughly $200M post-money in May 2026, co-led by Wing VC and Madrona with Snowflake Ventures, Hudson River Trading and Samsung Next (Gray Swan; Forbes Australia). The crowd is paid in prize money; the company is valued on software.

This is the only vertical in the atlas where the buyer's money is new, the supply is genuinely scarce, and the spread is not capped by what a rival can pay for an hour of labour. It is also the one where the evidence is thinnest — nobody in it discloses revenue.

Whose budget

Two pools, both created after 2023, neither displacing an incumbent.

The first is frontier labs. Pre-release adversarial evaluation is now compulsory by norm — Gray Swan says its work is cited in 11 frontier model system cards including GPT-5, with Anthropic, OpenAI, Meta and the UK AI Security Institute as verified customers and Google DeepMind, xAI, Amazon, Snowflake, ByteDance, ElevenLabs, Intercom and Deloitte claimed (Forbes Australia; RL List). Irregular's work appears in the Claude 3.7 Sonnet and OpenAI o3/o4-mini system cards (TechCrunch).

The second is enterprises deploying agents under the EU AI Act, NIST AI RMF and SOC 2 — slower, because it runs through procurement, but recurring.

What matters more than the size is the shape. Nobody was going to hire a head of adversarial evaluation in 2023, so there is no salary the price anchors to. Compare SDR-as-a-service, where every retainer is measured against ~$100k of loaded US SDR cost. That is a displacement sale; this is not. Hence budget: 5 — the buyer is a lab signing against a launch date, not a finance team comparing headcount.

Published pricing, from a single vendor blog and therefore weak: $8k–$15k for a simple chatbot, $15k–$35k for a RAG pipeline, $25k–$60k for a tool-using agent, $50k–$150k+ for multi-agent and MCP ecosystems, or $5k–$20k/month continuous. The same source claims a $2.26B market in 2026 (aivyuh) [WEAK — vendor blog, no methodology].

Can you get the supply

This is the strongest supply-side position in the sweep, and the reason is social rather than logistical.

Elite jailbreakers are a small, self-selecting, competitive population. You cannot post a job advert for them; money alone does not assemble them, because the people who are good at this are motivated by leaderboards and reputation at least as much as by cash. The prize mechanic is the recruiting instrument. Gray Swan's ~15,000 Arena red-teamers are an asset in the way that a resume database is not.

Contrast the general bounty platforms, which have far larger crowds and far less pricing power: HackerOne at ~300,000 researchers, Bugcrowd at ~500,000, Synack at ~1,500 vetted SRT members (Ciphers Security). Synack's small vetted pool commands $100k–$300k+/year flat-credit contracts; the large open pools sell VDP subscriptions from $10k. Scarcity of the crowd, not size of it, is what prices. supply: 5 — hard is good, and this is the hardest.

That 5 survives the test the spread score below fails, and it is worth saying why. The supply claim is not an inference from an undisclosed number: Synack's ~1,500 vetted researchers command $100k–$300k+/year contracts while HackerOne's ~300,000 and Bugcrowd's ~500,000 sell subscriptions from $10k — a 200-fold difference in crowd size running the opposite way to price. That is an observed funnel, which is what the rubric asks you to score. The honest qualification is the framework's warning about adjacent assets: every lab runs its own bounty page and could convene this population directly. That risk is priced below in hold: 3, not here — the labs' reach does not make the crowd easier for a new entrant to assemble.

speed: 4. A challenge can be live in a fortnight: there is no rig to buy, no licence to obtain and no roster to train, and the prize pool is funded by the buyer rather than out of your working capital — the UK AISI challenge put $171,800 in front of the crowd on the buyer's money. What holds it off a 5 is the Which side you build first shape named at the bottom of this page: a crowd is worthless without challenges and a challenge is worthless without a lab willing to expose a pre-release model, and that first exposure is a trust relationship with a safety team, not a purchase order a founder can sign. Weeks once you have the first lab; a quarter or more to get it.

What the spread looks like

Forbes reports that Gray Swan's enterprise revenue comes from Shade (automated red-teaming) and Cygnal (runtime guardrails) — its software products — not from Arena, the crowd platform. Arena's output trains the products (Forbes Australia).

Take rate is inferred, not sourced

Gray Swan discloses no revenue. The ">90% effective take" in this page's frontmatter is an inference from two facts that are separately sourced — challenge pools of $40k–$300k+ paid to the crowd, and enterprise revenue attributed to software rather than to Arena — and nothing bridges them. No filing, leak or press report gives Gray Swan's revenue or its cost of crowd. Treat the number as a structural argument about where value accrues, not as a measurement. It is [UNVERIFIED] and will stay that way until somebody sees the accounts.

What is documented is how little the individual earns. A profiled top red-teamer made $10,000 across 1,000+ challenges in roughly a year (MBG). Lab bounties are similar: OpenAI pays up to $100k but standard awards run $200–$20,000, Anthropic caps at $15,000, Google's AI VRP starts at $30,000 — and median payouts across programmes sit at $500–$5,000 (Wraith).

Synack states the mechanism openly: buyers purchase flat credits and "researchers are compensated internally by Synack" (Ciphers Security). The platform absorbs the variance and keeps the spread — the same structure as Expert networks, executed on a population that accepts prize money instead of an hourly rate.

spread: 4, revised down from 5. The reason is the flag directly above, applied honestly. This page's own text says the >90% take "is a structural argument about where value accrues, not a measurement," and the rubric is explicit that a 4 on a guess is not a 4 — so a 5 on a guess certainly is not a 5. Compare what the 5 in that rubric is anchored to: Expert networks pays the expert $200/hour and bills the client $800–$900, both sides published, across four decades. Here one side of the trade is documented in unusual detail — $171,800 in prize pools, $10,000 a year to a top performer, median payouts of $500–$5,000 — and the other side does not exist in the record at all. Gray Swan discloses no revenue; neither does Irregular, Haize or Patronus; the vertical has no observed cash valuation at any point in its history.

That is exactly the evidence standard Voice, speech and low-resource language data carries at spread: 4 — a richly sourced supply side multiplied by an assumed buy side — and ranking the two differently would mean paying this page for a number nobody has seen. It stays at 4 rather than 3 because the level, if the structure holds, is genuinely at the top of the atlas and half the arithmetic is real. It cannot be a 5 until somebody sees the accounts. See What a rake can actually be.

Can you hold it

Weaker than the rest of the scorecard, and this is where the model is exposed.

The crowd can leave. Every frontier lab runs its own bounty programme, and a researcher who wins a Gray Swan challenge learns exactly which lab's bounty page to visit next. That is textbook Getting cut out. What holds them is not a contract but the fact that a leaderboard across many models pays better in expectation than any single programme.

The buyer side is stickier. Once your name is in eleven system cards, the next lab buys you because regulators and journalists recognise the name. That reputational lock-in compounds — and it is winner-take-most, which is bad news for entrant number four.

One customer is a binary event is the unmeasured risk: with 20+ enterprise customers claimed and a handful of labs doing frontier work, the buyer list is short, and nobody has published a concentration figure for any company here. hold: 3.

What AI does to it

The only vertical in the atlas scoring ai: 5 — capable models create this budget rather than eating it.

Every capability increase raises the cost of a failure, widens the attack surface (tool use, multi-agent, MCP), and adds a system card needing external evidence. The pricing ladder is explicitly capability-indexed: a chatbot is $8k–$15k, a multi-agent MCP ecosystem is $50k–$150k+. Models make their own audits more expensive.

The crowd is not safe from the same force. Shade is automated red-teaming — Gray Swan uses the crowd's output to build the thing that reduces its dependence on the crowd. That is either the smartest move in the sector or the mechanism by which the crowd notices it is training its replacement. Both readings are live; see What better models do to each layer.

The 5 holds where the spread score did not, on the same standard. This score asks whether models grow the budget or delete it, and the answer is visible from outside the companies: 11 frontier model system cards including GPT-5, none of which existed in 2023, against a pricing ladder indexed to capability. Both are facts about the buyer's behaviour rather than about undisclosed accounts, so nothing here rests on the inference that sank the spread. The Shade risk is a claim about who captures the budget, not about whether it grows — and it is scored in spread and hold, where it has already cost this page a point.

What would kill it

What would kill it

Revenue never arrives. The Information published "Revenue Lags at AI Evaluation Startups" on 14 April 2025 — the one trade-press piece squarely questioning whether the marks in this slice are supported. The atlas has the headline and dateline only; the article itself is paywalled and was not read, so treat it as a signal rather than as evidence. Haize Labs reached ~$100M valuation after seven months on a General Catalyst seed with no disclosed revenue (PitchBook); no 2025–26 news has surfaced since.

Three other ways it ends. Labs in-house it — the pattern documented across Expert data for frontier labs, where labs keep what is strategic and outsource what is spiky; adversarial evaluation is arguably strategic. Consolidation caps the prize — Lakera went to Check Point in 2025 (Seeking Alpha), Robust Intelligence to Cisco in 2024; a vertical whose exits are strategic tuck-ins does not support $200M marks at scale (What the public market pays for labour). The crowd revolts — $10,000 a year for a top performer on a platform valued at $200M is the arithmetic that gets posted to a forum.

Who is already there

CompanyPositionMoneyEvidence quality
Gray Swan AIArena crowd (~15k), Shade, Cygnal$40M Series A, ~$200M post, May 2026Funding solid; revenue undisclosed
Irregular (ex-Pattern Labs)Frontier lab evals$80M at $450M post, Sequoia + RedpointFunding solid; revenue undisclosed
Haize LabsAutomated red-teaming~$100M valuation, GC seed Aug 2024No 2025–26 coverage found
Patronus AIEval tooling, agent stress-testing$50M, Jun 2026; ~$70M+ totalRevenue undisclosed
LakeraAcquired by Check Point, 2025Terms undisclosedExit price unknown
HackerOne / Bugcrowd / SynackGeneral bounty platformsSubscription + payout feesPricing published

Adjacent but distinct: Prolific pivoted into evals, RLHF and red-teaming, reaching an estimated $350M annualised by April 2026, and is the only company in the wider sector that publishes its take rate — 42.8% (Prolific). The best proxy for what a disclosed spread on adversarial human work looks like.

Where the record is thin

What nobody has published

No revenue or customer-concentration figure exists for Gray Swan, Irregular, Haize or Patronus. The only pricing source for enterprise engagements is a single vendor blog with no stated methodology, and the $2.26B market size comes from that same page. Gray Swan's split between Arena costs and Shade/Cygnal revenue is asserted by one Forbes piece and confirmed nowhere. Lakera's and Robust Intelligence's exit prices were never disclosed, so the vertical has no observed cash valuation at any point in its history.

Two further holes. Nobody has measured how much of the crowd is shared with the labs' own bounty programmes — that Getting cut out rate decides whether the >90% take is durable or temporary. And the Which side you build first problem is unusually specific: a red-team crowd is worthless without challenges, and challenges are worthless without a lab willing to expose a pre-release model. Gray Swan solved it with prize pools funded by the buyers themselves. Whether a second entrant can is untested.