Capability Gap

All companies

Gray Swan AI

15,000 red-teamers paid in prize money; revenue from the software their attacks train. The widest inferred spread in the atlas, on the thinnest evidence.

low confidence4 minupdated 2026-08-29red-teaming · evals · crowds · ai safety
Vertical
Evals and red-teaming
Founded
Not established in the record
Headquarters
Not established in the record
Raised
~$45M+
Last valuation
~$200M post-money (May 2026)
Revenue
Not disclosed
Status
Active — cited in 11 frontier model system cards

Gray Swan runs three products off one asset. Arena is a crowd of roughly 15,000 red-teamers who attack models in timed, prize-funded challenges. Shade is automated red-teaming. Cygnal is runtime guardrails. The crowd generates attack data; the software is what enterprises buy (Forbes Australia; RL List).

That separation is the company. Forbes states it directly: enterprise revenue comes from Shade and Cygnal, not from Arena — Arena's output trains the products. The crowd is compensated in challenge prizes; the value accrues to a software line.

The money

$40M Series A at roughly $200M post-money, May 2026, co-led by Wing VC and Madrona with Snowflake Ventures, Hudson River Trading and Samsung Next (Gray Swan). Total raised is approximately $45M+; the atlas register carries no earlier priced round.

Revenue: not disclosed, anywhere, at any point. Neither is customer count beyond "20+ enterprise customers", nor gross margin, nor the cost of the crowd. Every economic claim on this page is an inference.

Customers

Verified: Anthropic, OpenAI, Meta, the UK AI Security Institute. Claimed but not independently confirmed: Google DeepMind, xAI, Amazon, Snowflake, ByteDance, ElevenLabs, Intercom, Deloitte (Forbes Australia; RL List).

The distribution asset is not the customer list, it is the citation list: Gray Swan's work appears in 11 frontier model system cards, including GPT-5. A system card mention is a durable, public, third-party endorsement that a competitor cannot buy. It is the closest thing to a moat on the buyer side of Adversarial evals and red-team crowds.

Note the shape of that customer base against One customer is a binary event. A handful of frontier labs plus 20-odd enterprises is a short list, and nobody has published what share the largest one represents. Appen was worth $4.3B with 80% of revenue in five clients before Google left.

What the crowd is paid

This is the documented half of the business.

The UK AISI challenge run with OpenAI, Anthropic and Google DeepMind carried a total prize pool of $171,800: $86,000 on the leaderboard, $45,000 distributed by volume, $25,800 in first-break bounties at $300 each, $10,000 for over-refusal, $5,000 for new users, with a $100 minimum payout threshold. Gray Swan challenge pools generally range $40,000–$300,000+ (MBG Security).

At the individual level: a profiled top red-teamer earned $10,000 across 1,000+ challenges in roughly a year.

Set that against the alternative on offer. Lab-run bounties pay OpenAI up to $100k (standard awards $200–$20,000), Anthropic max $15,000, Google's AI VRP from $30,000 — with median payouts of $500–$5,000 across programmes (Wraith). Arena competes for the same people, and the people can go direct. That is the Getting cut out risk, and it is unquantified.

The take rate, and why it is a guess

Inferred, not sourced

The atlas estimates Gray Swan's effective take on its crowd's output at >90% [UNVERIFIED]. That number is built from two separately sourced facts — prize pools in the tens to low hundreds of thousands, and enterprise revenue attributed to Shade and Cygnal rather than Arena — with nothing connecting them. Gray Swan has never disclosed revenue or cost of crowd. The estimate is a structural argument about where value lands, not a measurement, and it should not be typed into a model. Compare Expert networks, where the ~70–80% take is observed from both sides of the trade, or Prolific, which publishes a 42.8% platform fee on its own pricing page. See What a rake can actually be and GMV is not revenue.

The one comparable that states its mechanism openly is Synack: buyers purchase flat credits and "researchers are compensated internally by Synack". The platform absorbs the variance and keeps the difference. Gray Swan's structure looks like that with a leaderboard bolted on and a software business attached to the exhaust.

What to watch

Whether Shade cannibalises Arena. Automated red-teaming trained on crowd output is, eventually, a substitute for crowd output. That is either the correct strategy — convert transient human labour into a renewing asset, the same move Distyl AI is attempting in Forward-deployed engineering — or the moment the crowd realises what it is building. Both readings are live; see What better models do to each layer.

Whether the crowd stays. $10,000 a year for a top performer, on a platform marked at $200M, is the arithmetic that gets posted to a forum. The prize mechanic works because status is part of the compensation; status is fragile.

Whether revenue arrives. The Information published "Revenue Lags at AI Evaluation Startups" on 14 April 2025 — the atlas has the headline and dateline, not the article. Haize Labs reached a ~$100M valuation seven months after founding on a General Catalyst seed (PitchBook) and has generated no coverage since. Irregular raised $80M at $450M post from Sequoia and Redpoint (TechCrunch). None of these companies has published a revenue figure; the vertical has no observed cash valuation in its history — Lakera's sale to Check Point and Robust Intelligence's to Cisco were both undisclosed. See What the public market pays for labour.

What is missing

Founding date, headquarters, headcount, revenue, gross margin, customer concentration, cost of crowd, and the Arena/Shade/Cygnal revenue split. The single source for the crowd-versus-software separation is one Forbes piece. For a company whose valuation rests entirely on that separation being true, the record is close to empty.