Outlier AI advertises up to $40/hour for English work and about $7.50/hour for Indian-language projects (remowork). The buyer pays more for the rarer language, not less. That inversion — pay less to acquire the scarcer input, charge more to sell it — is the entire business.
It is also, as far as the public record goes, unmeasured. This vertical has the best supply-side data in the atlas and the worst buy-side data. The reason is structural rather than sinister: contributor rates are advertised because recruiting contributors is marketing, while Defined.ai, Shaip and LXT route every buyer through "get in touch". You can find out to the cent what a person is paid to read a paragraph in Yoruba. You cannot find out what anyone charges for the result.
That is why this page carries confidence: low and why its take rate is stated as an estimate rather than a finding.
Whose budget
Moderate and bifurcating, which is the polite way of saying shrinking at one end.
Multilingual TTS/ASR teams and voice-agent builders — ElevenLabs, Cartesia, Deepgram, plus every lab's audio group — still buy spec'd speech. But synthetic speech and self-supervised pretraining have eaten a great deal of the mid-market. The budget type is mixed: at a lab it is a new line inside an existing data budget; at an enterprise building a voice agent it displaces nothing and has to be argued from scratch.
budget: 2. This is the lowest budget score in the labour section, and it is a comparative judgement, not an absolute one. Set it against Robotics teleoperation and physical-world data, where humanoid programmes are capitalised at $100M+ with data named in the plan, or Adversarial evals and red-team crowds, where a launch date forces the purchase. Speech data is a mature input that the buyer has been acquiring for a decade and increasingly knows how to generate.
Not one major vendor publishes a rate card. Shaip's own costing guide lists pricing units — per second, per minute, per produced hour — and no numbers (Shaip). Defined.ai lists 819+ datasets with no prices, all behind an enquiry form (Defined.ai). Troveo states the position plainly: "Treat public rate cards as starting points; most deals at scale are negotiated… Custom collection is priced per produced hour and rises with language rarity and specification detail" (Troveo). The $80–$300+ per produced hour figure used later on this page is a working assumption derived from supply-side rates plus normal data-vendor margins. It is [WEAK]. Treat it as a hypothesis to be tested with a quote request, not as a market price.
Can you get the supply
Two populations with nothing in common except the word "voice".
Major languages: no. English and the top twenty languages are trivially commoditised, and the advertised rates prove it — you can buy an hour of recorded English conversational speech from a dozen platforms tomorrow.
Genuinely low-resource languages: yes. Yoruba, Pashto, Quechua, dialectal Arabic. You need on-the-ground recruiters, local payment rails, literacy-appropriate protocols and consent documentation that will survive a lawyer. That is a logistics business in the countries concerned, not a marketplace — see Where the supply can legally live.
There is a third scarcity, and it was created by litigation rather than logistics: rights-cleared, provably-consented voice. Troveo sells on the line "speakers opted in and get paid"; Luel sells a rights-cleared marketplace. Consent provenance is a legal asset, and legal assets do not commoditise the way logistical ones do.
supply: 3, blended. Score the English tier at 1 and the low-resource consented tier at 4.
speed: 3, blended the same way. On the English tier you can be recording next week — the platforms in the pay table above recruit contributors with an advertised hourly rate and nothing else. The tier worth being in is slower: on-the-ground recruiters, local payment rails in the countries concerned, literacy-appropriate protocols and consent paperwork that will survive a lawyer are a per-country build measured in months, not weeks. The buy side adds its own delay, because there is no rate card to transact against — Defined.ai, Shaip and LXT route every buyer through an enquiry form and negotiate per produced hour, so the first invoice waits on a bespoke quote and a delivered corpus rather than on a subscription anyone can sign. Months to a real invoice on the defensible half of the vertical.
What the spread looks like
The supply side, in the industry's own recruiting copy:
| Platform | Contributor pay |
|---|---|
| Babel Audio | $17.50–$150 per recorded hour; $50/recorded hr for video conversations |
| Lionbridge | $40–$150 per task |
| TELUS Digital | $14–$35/hr (annotation / transcription) |
| Outlier AI | up to $40/hr English; ~$7.50/hr Indian-language projects |
| Luel | ~$0.50/min approved ≈ $30/recorded hr; $15/hr bilingual conversations |
| DataAnnotation.tech | $20–$30+/hr audio eval and voice acting |
| CrowdGen (Appen) | $10–$20/hr, flat per-recording fees varying by country |
| Innodata | $15–$20/hr |
| Pila8 | ~$6–$30/hr via credits |
| Silencio | $10+/hr, paid in USDC |
| CloudFactory | "local-market rates" — Nepal, Kenya, Philippines, Colombia |
Source: remowork.
Now the arithmetic that cannot be completed. If a buyer pays $100–$300 per produced hour for spec'd low-resource speech and a contributor gets $10–$30 per recorded hour — and produced hours are always fewer than recorded hours once rejection, QA and transcription are taken out — the take lands somewhere around 60–85%.
The supply figure is well sourced. The buy figure is not sourced at all. Multiplying a documented number by an assumed one produces an assumed number. spread: 4 in the frontmatter reflects where this vertical would sit if the assumption holds; it is not a measurement, and no reader should carry it into a model. Compare Expert networks, where both sides of the trade are published and the ~70–80% take is a fact, or Prolific, which prints its 42.8% platform fee on its pricing page. See What a rake can actually be.
Can you hold it
Weakly. hold: 2.
Datasets are delivered, not subscribed. Once a buyer has the corpus, the buyer has the corpus — the transaction ends and the relationship has to be re-won on the next specification. That is the opposite of Expert networks, where the buyer needs a new expert every week and therefore needs the network permanently.
Contributors have no reason to stay either. The pay table above is the pay table every competitor can read, and the same people appear on several platforms at once. There is no leaderboard, no reputation asset, no rig. The Getting cut out risk is not that buyer and seller meet directly; it is that neither side has a reason to prefer you.
The exception is consent provenance. A buyer who has built a product on a corpus with documented, contract-backed consent has switching costs measured in legal risk. That is worth holding onto and is the single strongest argument for being in this vertical at all.
What AI does to it
ai: 2 — the second-lowest score in the atlas's labour section, and the honest one.
Synthetic speech is good and getting better, and self-supervised pretraining removed a large share of the labelled-audio requirement that funded this industry. Appen, the only public pure-play, shows what that looks like in audited accounts: operating revenue of $230.8M in FY2025, with Appen Global down 21.1% to $127.9M, gross margin 40.3%, and crowd NPS falling from 33 to 22 (Appen FY2025 Annual Report). Its China segment grew 74.8%; its legacy business did not.
What models cannot do is manufacture a consented recording of a real speaker of a language with two million speakers and no written corpus. That residue is real, and it is small. See What better models do to each layer.
What would kill it
Commodity English speech is already gone, on the same curve that took undifferentiated egocentric video to free in Robotics teleoperation and physical-world data and legal document review from $460k to $36k per matter. The mechanism is identical: when the unit of work becomes model-tractable, price falls 70–95% in about eighteen months. There is no reason to believe major-language speech is exempt, and Appen Global's -21.1% is what the early part of that curve looks like on a balance sheet.
Two further ways it ends. The buyer generates it — a lab with a good TTS model can synthesise training data for its own ASR, closing the loop entirely. The consent premium gets regulated into a commodity — if provenance documentation becomes a standardised, auditable checklist rather than a bespoke legal asset, the scarcity Troveo and Luel are selling turns into a form anyone can fill in.
Who is already there
Legacy, and mostly struggling: Appen (1M+ contributors, structural decline since losing big-tech contracts), Defined.ai (819+ datasets, no published prices), Shaip, LXT, Welocalize, TELUS Digital (ex-Lionbridge AI), Innodata, CloudFactory.
Newer, and selling provenance rather than volume: Troveo (rights-cleared, contributors opted in and paid), Babel Audio, Pila8, Silencio Voice AI (pays in USDC), Luel.
The newcomers compete on a legal property; the incumbents compete on scale, and scale just stopped being scarce. Innodata, which does this work among others, posted Q2 2026 revenue of $92.1M at +58% YoY with 49% adjusted gross margin (StockTitan) — a diversified BPO can still compound on this demand even as the pure speech tier deflates.
Where the record is thin
Almost everywhere on the demand side.
Supply-side rates are richly documented because contributor recruitment is marketing — every platform in the table above wants those numbers found. Buy-side pricing is documented nowhere, because a negotiated per-produced-hour price is the vendor's only leverage. The result is that the most confident-sounding number in this vertical, the take rate, is the one with no evidence behind it. Defined.ai, Shaip and LXT publish no buy-side price at all; none of the three discloses revenue; no gross margin figure exists for any private company in this vertical.
What would settle it: three quotes for the same specification — say 500 produced hours of consented conversational Yoruba — from Defined.ai, Shaip and a newcomer like Troveo, set against the advertised contributor rate for the same work. That is a fortnight of procurement effort and it would convert this page from confidence: low to something worth acting on. Until then, the structural question — whether the winner here is a marketplace or a managed collection agency — cannot be answered, because the margin that would decide it has never been seen.