Audio & voice
Turn a script into a natural voiceover
I have a 2,000-word script. Which text-to-speech actually sounds human, and what will it really cost me?
Last checked 2026-08-29How we ranked this ↓
1
ElevenLabs
ElevenLabsStill the only one built around a whole script, not a request.
- Use it when
- You have a real script — an explainer, an audiobook chapter, a course module — and you need pronunciation fixed, two voices swapping mid-paragraph, and a WAV you can hand to an editor without writing chunking code.
- Cost
- Free 10k credits/mo, no commercial use. Paid $6-$990/mo (30k-6M credits).Free $0 (10k credits/mo), Starter $6 (30k), Creator $22 (121k, first month half price at $11), Pro $99 (600k), Scale $299 (1.8M), Business $990 (6M), Enterprise custom. Unused credits roll over but only up to two months' worth, and only while the subscription stays active. Instant voice cloning unlocks at Starter; professional voice cloning at Creator. Pay-as-you-go top-up credits last 12 months and kick in when the monthly quota runs dry.
- Good at
- Studio is the actual differentiator: import EPUB, PDF, DOCX, TXT, HTML or a URL, split into up to 500 chapters, assign different voices inside a single paragraph, apply pronunciation dictionaries with phoneme or alias tags, export MP3, WAV or video. Nothing else here treats a script as a document.
- The catch
- It is far and away the most expensive per character on this list — the mid plans work out around ten to twenty times what Speechify or Inworld charge per million characters. And the free tier is a trap dressed as generosity: ElevenLabs states plainly that it includes no commercial licence and cannot be used for any commercial purpose, and that free-plan output published anywhere must carry 'elevenlabs.io' attribution in the title. You need a paid plan before a single second of this reaches a client.
- Wrong for
- Anyone generating narration at volume, or anyone who just needs an API call. Go to Speechify or Inworld and keep the difference.
2
Cartesia Sonic 3.6
CartesiaTop of the blind arena. Bring your own script pipeline.
- Use it when
- You are generating voiceover from code, naturalness is the thing you will be judged on, and you are happy to write your own paragraph splitting and stitching.
- Cost
- Free 20k credits/mo, no commercial use. Paid $5-$299/mo (100k-8M credits).Free $0 (20k credits, ~27 TTS minutes), Pro $5 (100k credits, ~133 min), Startup $49 (1.25M credits, ~1,667 min), Scale $299 (8M credits, ~10,667 min), Enterprise custom. Commercial licence starts at Pro. Instant voice cloning from Pro (2 clones); professional cloning from Startup. Artificial Analysis lists Sonic 3.6 at $49 per 1M characters.
- Good at
- Raw naturalness. It ranks first on both Artificial Analysis speech arenas at Elo 1288, with the gap to third wider than the gap from third to twelfth. 61 locales, 500+ preset voices, Hinglish code-switching, and accent adherence that survives voice cloning. Sub-90ms replies and ~132 characters/sec generation, so a long script renders fast.
- The catch
- It is engineered for live voice agents, not narration. The launch material sells sub-90ms latency and a 99.9% uptime SLA — there is no Studio, no chapters, no document import, no pronunciation dictionary workflow. You will spend day three writing the splitter, the retry logic and the concatenation that ElevenLabs gives you in a text box. The free tier's 20k credits is roughly 27 minutes and carries no commercial rights.
- Wrong for
- Anyone who wants to paste a script into a browser and download an MP3. That is ElevenLabs or Speechify.
3
Speechify (Simba 3.2 API)
SpeechifyNear arena-leading quality at a tenth of ElevenLabs' price.
- Use it when
- You are shipping narration at volume and the budget matters more than the last five percent of naturalness — or you want a free tier you can legally publish from.
- Cost
- Free 50k chars/mo, commercial use allowed. Paid $10-$499/mo, $6-$10 per extra 1M chars.Free $0 (50k TTS characters/mo, hard cap, no overages, commercial use permitted). Starter $10/mo then $10 per additional 1M characters. Pro $99/mo then $8 per 1M. Scale $499/mo then $6 per 1M. Enterprise volume pricing. Artificial Analysis lists Simba 3.2 at $10 per 1M characters at Elo 1243 — second on the leaderboard. 1,500+ voices, 30+ languages.
- Good at
- Price-to-quality, by a distance. Simba 3.2 sits 45 Elo behind the arena leader and costs roughly a fifth per character. The free tier is the only one here that explicitly permits commercial use, and it is a hard cap — no overage, no surprise invoice at the end of the month.
- The catch
- Voice cloning is not on the free tier. The Free card lists commercial use and 1,500+ voices; 'Voice cloning, streaming, SSML' appears only from Starter. So the free 50k characters gets you stock voices only — if the brief was 'clone the founder's voice', you are paying from day one. The free tier's hard cap also means generation simply stops mid-project rather than degrading.
- Wrong for
- Long documents needing chapter structure and pronunciation dictionaries. Use ElevenLabs Studio for that.
4
Gemini TTS
GoogleFree two-speaker dialogue in 100+ languages, with a data catch.
- Use it when
- You need a two-hander — interview, dialogue, back-and-forth explainer — or you are prototyping and want to spend nothing on non-confidential text.
- Cost
- Free tier free of charge. Paid $1/1M text-in, $20/1M audio-out tokens.gemini-3.1-flash-tts-preview: paid tier $1.00 per 1M input text tokens, $20.00 per 1M output audio tokens; batch halves both to $0.50 and $10.00. Free tier is listed as free of charge for both input and output. Free-tier rate limits are no longer enumerated in the docs — they are only visible in AI Studio under your own account. Also available: Gemini 2.5 Flash Preview TTS and 2.5 Pro Preview TTS.
- Good at
- Native multi-speaker generation — up to 2 speakers in one call, which no other tool here does without stitching separate renders. Over 100 languages with automatic detection, a 32k-token session context, and genuinely free access for prototyping.
- The catch
- Two problems land on day three. Google's own pricing page says free-tier content is 'used to improve our products' while paid-tier content is not — so a client's unreleased script cannot go through the free path. And the speech-generation docs warn that quality and consistency 'may begin to drift with generated outputs that are longer than a few minutes', which is exactly the length of the voiceover you are trying to make. Add: 30 preset voices, no voice cloning, and every TTS model still carries a 'preview' label.
- Wrong for
- Anyone who needs a specific or cloned voice, or a 20-minute continuous read. Go to ElevenLabs or Speechify.
5
Chatterbox Multilingual v3
Resemble AIMIT weights, free cloning, no per-character meter running.
- Use it when
- Volume is high enough that per-character pricing hurts, the script cannot leave your infrastructure, or you want unlimited voice cloning without a subscription tier deciding when.
- Cost
- Free. MIT-licensed weights; you pay only for the GPU.No pricing. Weights are MIT — the genuinely permissive licence in this field, and the reason it beats Fish Audio S2 Pro and Breeze TTS 2 here despite both scoring higher on the arena. 500M parameters. Runs on CUDA, CPU or Apple Silicon (mps); the Nano variant runs 3x faster than realtime on 8 CPU cores. 23+ languages plus dedicated single-language packs for six.
- Good at
- Zero-shot voice cloning from roughly a 10-second reference clip, with no tier gate and no per-clone limit. Exaggeration and cfg_weight parameters give you direct control over how dramatic the read is, and Turbo/Nano support paralinguistic tags like [laugh] and [cough].
- The catch
- Every file it generates carries Resemble's Perth neural watermark — imperceptible, but designed to survive MP3 compression and editing, and not something you can switch off. On top of that, output quality depends on you tuning cfg_weight and exaggeration per reference clip, and a reference whose language does not match the language tag will bleed its accent into the read. That is an afternoon of fiddling per voice, every voice.
- Wrong for
- Anyone without a GPU and an afternoon. If you just need a clean read by Friday, pay Speechify $10.
Also considered
What we left out, and why. A list is only trustworthy if you can see what it rejected.
- Kokoro-82M — Apache 2.0 and only 82M parameters, so it runs anywhere — but Elo 1059 on the arena and no voice cloning at all. Chatterbox does the same job better under an equally free licence.
- Fish Audio S2 Pro — Open weights in name only. The Fish Audio Research License forbids any commercial purpose without a separate written agreement, and requires 'Built with Fish Audio' displayed.
- Breeze TTS 2 — Leading open-weights model on the arena at Elo 1220, but the code is Apache 2.0 while the weights sit under a research-and-non-commercial licence requiring written permission from Resonia, Inc. Free to download, not free to sell.
- Inworld Realtime TTS-2 Flash — Genuinely cheap — $15 per 1M characters on-demand dropping to $7 on Growth, with ~70 TTS minutes free — and Elo 1228. Narrowly out on script-handling workflow rather than on price or quality.
- OpenAI TTS — As of 29 Aug 2026 the OpenAI API pricing page lists realtime and transcription audio models but no standalone text-to-speech model or per-character rate. Cannot verify a price, so cannot rank it.
- Qwen-Audio-3.0-TTS-Plus — Tied for second on the arena at Elo 1243, but $27.6 per 1M characters — nearly triple Speechify for the same score.
- Murf — Pricing page is JavaScript-only and returned no readable tiers on 29 Aug 2026. No verified price, no ranking.
What would change this list
- Whether ElevenLabs cuts per-character pricing. It is currently defending a roughly 10-20x premium on workflow alone, and Speechify at $10/1M with Elo 1243 is the pressure.
- Cartesia shipping a long-form Studio-equivalent. Sonic 3.6 already leads the arena; give it document import and chapters and it takes rank 1 outright.
- Any arena-competitive model landing under a real Apache or MIT licence. Breeze TTS 2 and Fish S2 Pro both beat Chatterbox on quality and both forbid commercial use — that stalemate breaking would reorder the bottom half.
- Gemini TTS leaving preview, and whether the paid tier fixes the multi-minute quality drift. If it does, free two-speaker narration becomes very hard to argue with.
- OpenAI reinstating a listed standalone TTS model and price. Its absence from the pricing page is the single biggest unknown here.
- Watermarking becoming mandatory. Chatterbox already embeds Perth; if hosted vendors follow, the self-host escape hatch closes.
How we ranked this
Sources
ElevenLabs pricingElevenLabs API pricingElevenLabs billing docs (credit rollover)ElevenLabs: can I publish the content I generate?ElevenLabs Studio product guideCartesia pricingCartesia: Introducing Sonic-3.6Artificial Analysis text-to-speech leaderboardTTS Arena v2 (blind A/B)Speechify pricingGemini API pricingGemini speech generation docsChatterbox on GitHubChatterbox model card (MIT, watermarking)Kokoro-82M model cardFish Audio S2 Pro licenceBreeze TTS 2 model card (non-commercial weights)Inworld pricingOpenAI API pricing