THE AI RANKINGS

development

Best Text-to-Speech APIs

The best text-to-speech APIs in 2026 compared for developers — ElevenLabs, Cartesia, Deepgram Aura-2, Rime, Google Cloud TTS, OpenAI, Inworld, Hume, Azure and Amazon Polly, with time-to-first-audio latency, streaming protocols, per-character and per-token pricing, voice cloning, SSML control and decisive picks for real-time voice agents, high-volume synthesis and self-hosting.

Updated August 2026

Quick answer: For real-time voice agents in 2026, Cartesia Sonic 3.5 is the latency leader at roughly 82ms end-to-end (about 40ms on Sonic Turbo), with Deepgram Aura-2 (sub-200ms time-to-first-byte, $30 per million characters) the natural pick for teams already running Deepgram speech-to-text, and Rime the specialist for phone agents that need conversational prosody. For the most expressive produced speech, the ElevenLabs API and its Eleven v3 model lead on quality across 70+ languages, while its Flash v2.5 model handles the real-time tier at around 75ms. For the cheapest high-quality synthesis at scale, Google Cloud Gemini 3.5 Flash TTS costs about $6 per million output tokens — roughly $0.09 for a 10-minute narration. For emotion-critical delivery, Hume Octave 2 leads on explicit emotional control; for enterprise compliance and language breadth, Microsoft Azure AI Speech covers 140+ languages with a HIPAA and SOC story; for high-volume AWS-native workloads, Amazon Polly is the cost-predictable workhorse. The one thing to decide first: whether you need a standalone TTS API at all, because bundled speech-to-speech models such as OpenAI GPT-Realtime-2 now collapse the whole voice pipeline into one call for many agent builds.

This guide ranks the APIs you call programmatically to turn text into speech inside a product — measured on latency, streaming, price and control, not on a web dashboard. If you want a creator or consumer tool with a web interface, a voice library and produced voice-overs, that is a different job: see our best AI voice generator guide, which covers ElevenLabs, Murf, Speechify and the rest as tools rather than endpoints. For the inverse problem — turning speech into text — see best AI for transcription. For cloning-specific picks and the legal landscape around it, see best AI voice clone.


The current state of text-to-speech APIs: August 2026

The text-to-speech API market has stopped competing on raw naturalness and started competing on latency, price and control.

The market backing this is large and growing. The text-to-speech market sits near $4.36 billion in 2026, up from $3.87 billion in 2025, and is forecast to reach roughly $7.9 billion by 2031 at a compound annual growth rate around 12.7% (Mordor Intelligence). Forecasts for the adjacent AI voice generator category run considerably higher, so treat any single headline number as one analyst’s definition of the market rather than a settled figure. The growth driver is no longer audiobooks or e-learning voice-overs — it is real-time voice agents, where speech is generated live inside a phone call or assistant and every millisecond of delay is audible.

Five shifts define the current moment.

First, latency became the battleground. The fastest APIs are now measured in tens of milliseconds of time-to-first-audio: Cartesia Sonic Turbo reaches roughly 40ms, ElevenLabs Flash v2.5 around 75ms, and Deepgram Aura-2 hits a 90ms optimised figure with sub-200ms time-to-first-byte. Sub-300ms time-to-first-byte is the threshold below which turn-taking in a conversation feels natural, and the agent-focused APIs are all now under it.

Second, the market split into three layers. Expressive production models (ElevenLabs Eleven v3, Hume Octave 2, Inworld) optimise for quality and emotional range on non-real-time work; real-time agent models (Cartesia Sonic, Deepgram Aura-2, Rime Coda, ElevenLabs Flash) optimise for streaming latency; and high-volume cloud models (Google, Amazon Polly, Azure) optimise for cost and compliance at scale. Choosing well means choosing the layer before the vendor.

Third, billing moved from characters to tokens on the newest models. Traditional TTS APIs charge per character — $4 to $30 per million is the common band. The newest models from Google and OpenAI bill on audio tokens instead: Google Gemini 3.5 Flash TTS costs $6 per million output tokens, and OpenAI’s gpt-4o-mini-tts works out to roughly $15 per million characters at typical ratios. Token billing means you should estimate cost from audio duration, not text length.

Fourth, streaming and WebSocket delivery became standard. Every agent-grade API now supports chunked or WebSocket streaming so audio starts playing before the full sentence is synthesised. REST-only, file-return endpoints still exist for batch narration, but they are no longer competitive for conversational use.

Fifth, the speech-to-speech shift is reshaping whether you need a TTS API at all. OpenAI GPT-Realtime-2 and xAI Grok Voice bundle transcription, reasoning and speech into a single model, removing the separate text-to-speech step for some agent builds (TestingCatalog). A standalone TTS API still wins when you want a specific branded voice, cheaper high-volume synthesis, or a best-of-breed pipeline you control end to end.

One market exit sits behind all of this: PlayHT shut down permanently on 31 December 2025 after Meta acquihired the PlayAI team, deleting all accounts, voice clones and audio with no migration path. If you are choosing an API today, choose one with a durable business behind it.


The three types of text-to-speech API

Match the layer to the job before comparing vendors within it.

TypeWhat it optimises forExamplesBest when
Real-time agent APIStreaming latency and turn-takingCartesia Sonic, Deepgram Aura-2, Rime, ElevenLabs Flash, InworldYou are building a voice agent, phone system or live assistant
Expressive production APIVoice quality, emotion and cloningElevenLabs Eleven v3, Hume Octave 2, MiniMax SpeechYou are producing narration, characters or branded audio at high fidelity
High-volume cloud APICost, compliance and language breadthGoogle Cloud TTS, Amazon Polly, Microsoft AzureYou need cheap synthesis at scale inside an existing cloud
Self-hosted open-weightData control and zero per-character costKokoro-82M, Chatterbox, Qwen3-TTSYou need on-device, offline or data-sovereign speech

Most production voice systems in 2026 use more than one layer: a low-latency streaming model for the live conversational turn, and a cheaper cloud or self-hosted model for high-volume, latency-tolerant synthesis such as batch notifications.


Top text-to-speech APIs ranked (August 2026)

Ranked by overall suitability for a team building a new voice product today, weighting latency, voice quality, streaming support, price and language coverage. Latency is the vendor-reported time-to-first-audio or time-to-first-byte; verify against your own region and payload before committing.

RankAPITypeFlagship modelLatency (TTFA/TTFB)PriceBest for
1CartesiaReal-timeSonic 3.5 / Sonic Turbo~82ms / ~40ms~$35 / M charsLowest-latency voice agents
2ElevenLabs APIExpressive + real-timeEleven v3 / Flash v2.5~75ms (Flash)Credit-based, commercial from $5/moBest voice quality and cloning
3Deepgram Aura-2Real-timeAura-2~90ms, sub-200ms TTFB$30 / M charsVoice agents on the Deepgram stack
4Google Cloud TTSHigh-volumeGemini 3.5 Flash TTS / Chirp 3 HDModerate$6 / M tokens; $30 / M chars (Chirp 3 HD)Cheapest quality at scale
5RimeReal-timeCoda / Mist v2 / Arcanasub-200ms cloud, sub-100ms on-prem$30–50 / M charsPhone agents with conversational prosody
6OpenAIExpressive + S2Sgpt-4o-mini-tts / tts-1-hd / GPT-Realtime-2~160ms (Realtime-2)$15 / M chars (tts-1); ~$15 / M chars (mini-tts)OpenAI-native pipelines and speech-to-speech
7InworldReal-timeRealtime TTS-2sub-200msfrom ~$5 / M chars at scaleSteerable real-time voice with one billing account
8HumeExpressiveOctave 2 / EVIModerateusage-basedEmotion-critical delivery
9Microsoft Azure AI SpeechHigh-volumeNeural / Neural HDModerate$15–16 / M chars; $22 / M (HD)Enterprise compliance and language breadth
10Amazon PollyHigh-volumeNeural / GenerativeNear real-time$4–100 / M charsHigh-volume AWS-native synthesis

Also worth knowing: MiniMax Speech 2.8 serves 40+ languages with strong expressiveness at around $100 per million characters (HD), LMNT and Neuphonic are low-latency developer options priced at roughly $43.60 and $17.60 per million characters respectively, and for self-hosting the open-weight Kokoro-82M and Chatterbox models run locally at no per-character cost.


The best text-to-speech APIs compared

Ordered by current suitability for a new voice build, not by company size. Every price is list; verify against each vendor’s live pricing page before you commit spend.

1. Cartesia — lowest-latency voice agents

Type: Real-time Flagship: Sonic 3.5, with Sonic Turbo for the lowest latency Pricing: Around $35 per million characters, with a free tier Key features: State-space (SSM) architecture, WebSocket streaming, 3-second instant voice cloning, 40+ languages, consistent latency at P99

Cartesia is the latency leader for programmatic speech. Its Sonic 3.5 model reaches roughly 82ms end-to-end, and Sonic Turbo pushes model latency to around 40ms — fast enough that an agent can begin responding before the user finishes hearing their own last word. The state-space model architecture, rather than a standard transformer, keeps latency consistent even at the 99th percentile, which matters more for live conversation than a good median. Instant voice cloning works from a 3-second reference clip.

Why it wins: The lowest and most consistent streaming latency available, behind a clean developer API built for the “bring your own speech-to-text and language model” pattern that most custom voice agents use.

Limitations: Cartesia is built for developers, not content teams — there is no polished editing interface or large stock-voice catalogue. It competes on speed and API quality, not on a produced-content workflow.

Best for: Developers building real-time voice agents, IVR systems and live applications where turn-taking latency is the load-bearing requirement.


2. ElevenLabs API — best voice quality and cloning

Type: Expressive and real-time Flagship: Eleven v3 (produced) and Flash v2.5 (real-time) Pricing: Credit-based; commercial use unlocks from the $5/month Starter tier, with effective per-character cost varying by plan Key features: 70+ languages, 5,000+ voices, WebSocket streaming, audio tags for non-verbal sounds, instant and professional voice cloning

The ElevenLabs API is the quality benchmark. Its Eleven v3 model reached general availability on 14 March 2026 with audio tags, a 68% reduction in complex-text errors and support for 70+ languages, and it sits in the top tier of the Artificial Analysis Text-to-Speech leaderboard. The model split is the key API decision: Eleven v3 is the most expressive option but is not real-time, so for voice agents and live dialogue ElevenLabs directs developers to Flash v2.5, which delivers sub-75ms latency across 32 languages. Pick v3 for produced audio; pick Flash for conversation.

Why it wins: The best available voice quality and the deepest cloning options, behind an API with mature streaming, the largest voice library and broad language coverage.

Limitations: The credit-based billing draws the most developer complaints — failed generations and retries consume credits, so real costs often exceed the advertised per-character estimate. ElevenLabs is also a defendant in the BIPA class actions filed in Chicago federal court between 11 and 13 May 2026 over alleged unauthorised voiceprint training — though this is not an ElevenLabs-specific risk: the same coordinated set of nine suits names Adobe, Alphabet, Amazon, Apple, Meta, Microsoft, NVIDIA and Samsung, so Google Cloud TTS, Azure and Polly sit under the same cloud.

Best for: Products where voice quality directly affects revenue — branded assistants, characters, produced narration — and any team that needs both a production and a real-time model from one vendor.


3. Deepgram Aura-2 — best for voice agents on the Deepgram stack

Type: Real-time Flagship: Aura-2 Pricing: $30 per million characters ($0.030 per 1,000) pay-as-you-go, with a $200 free credit (Deepgram) Key features: 90ms optimised latency, sub-200ms time-to-first-byte, 40+ English voices and 10+ Spanish voices across 7 languages, tight pairing with Deepgram Nova speech-to-text

Deepgram built Aura-2 as the text-to-speech half of a voice-agent stack whose other half — Nova-3 and Flux speech-to-text — is already the developer default for transcription. Its defining advantage is latency measured inside the full agent loop: the time from language-model output to the first audio byte is very low, and Aura-2 stays within the sub-300ms time-to-first-byte threshold that natural conversation requires. For teams already transcribing with Deepgram, adding Aura-2 keeps speech-to-text and text-to-speech on one vendor, one SDK and one bill.

Why it wins: Enterprise-grade streaming latency at a flat, predictable $30 per million characters, with the tightest integration path for anyone already using Deepgram for transcription.

Limitations: Language coverage (7 languages) is narrower than the cloud providers, and the voice library, while large in English, is smaller than ElevenLabs. Aura-2 is aimed squarely at agents, not produced narration.

Best for: Real-time voice agents and call-centre systems, especially teams already running Deepgram Nova or Flux for speech-to-text.


4. Google Cloud TTS — cheapest quality at scale

Type: High-volume Flagship: Gemini 3.5 Flash TTS (token-billed) and Chirp 3 HD (character-billed) Pricing: Gemini 3.5 Flash TTS $6 per million output tokens; Chirp 3 HD $30 per million characters; Standard and WaveNet from $4 per million characters; a $300 new-account credit, plus a monthly free allowance of 4M characters on Standard and 1M on WaveNet — the Gemini TTS models are excluded from the free tier (Google) Key features: The most generous free allowance on the older voice tiers, token billing on the newest models, instant custom voice creation, Google Cloud identity and compliance

Google Cloud Text-to-Speech is the value and free-tier leader. The newer Gemini TTS models are token-billed and remarkably cheap: Gemini 3.5 Flash TTS costs $6 per million output tokens, which works out to roughly $0.09 for a 10-minute narration, and it sits in the Artificial Analysis top tier on quality. The older Chirp 3 HD voices are character-billed at $30 per million, add human disfluencies and emotional range, and support instant custom voice creation. The recurring monthly free allowance — 4M characters on Standard voices, 1M on WaveNet — makes Google the cheapest place to prototype a pipeline, with one caveat: the Gemini TTS models sit outside it, so free-tier testing runs on the older voices and the $300 new-account credit is what covers early Gemini TTS usage.

Why it wins: The lowest per-unit cost for top-tier quality, a genuinely usable recurring free allowance on the older voices, and native access to Google Cloud’s identity, logging and data-residency controls.

Limitations: Streaming latency is moderate rather than class-leading, so for the tightest real-time turn-taking a dedicated agent API like Cartesia or Deepgram will feel faster. Token billing on the Gemini models means you must estimate from audio duration, not character count.

Best for: Cost-sensitive high-volume synthesis, prototyping on a real free tier, and teams already building on Google Cloud.


5. Rime — best for phone agents with conversational prosody

Type: Real-time Flagship: Coda (recommended for new builds), with Mist v2 and Arcana Pricing: Per-model, roughly $30–50 per million characters (Mist ~$30, Arcana ~$40, Coda ~$50), with free minutes on the Starter plan (Rime) Key features: 300+ voices, sub-200ms cloud latency (sub-100ms on-prem), mid-conversation code-switching that preserves voice identity, on-premise deployment

Rime is a text-to-speech API built for one job: making enterprise phone voice agents sound like people rather than IVR menus. Its models are tuned for authentic conversational prosody — the hesitations, emphasis and rhythm of real speech — rather than the smooth, even delivery of narration engines. Rime targets developers with 300+ voices and sub-200ms cloud latency, dropping to sub-100ms with on-premise deployment for teams that need speech to stay inside their own network. Its newer Coda model is the recommended default for new applications, with Mist v2 for high-performance general use.

Why it wins: Conversational-first voice quality plus an on-premise option, aimed precisely at the high-volume phone-agent use case where prosody and data control both matter.

Limitations: Rime is a specialist — it is not the tool for produced audiobooks or multilingual narration, and its language breadth trails the cloud providers. Model lineups move quickly, so pin a specific model version in production.

Best for: Enterprise phone and contact-centre voice agents that need natural conversational delivery and, optionally, on-premise deployment.


6. OpenAI — best for OpenAI-native pipelines and speech-to-speech

Type: Expressive and speech-to-speech Flagship: gpt-4o-mini-tts, tts-1 / tts-1-hd, and GPT-Realtime-2 for speech-to-speech Pricing: tts-1 $15 and tts-1-hd $30 per million characters; gpt-4o-mini-tts token-billed at roughly $15 per million characters (about $0.015 per minute of audio); GPT-Realtime-2 $32/$64 per million audio tokens in/out Key features: Streaming output, GPT-5-class reasoning in the Realtime line, one account for text, speech and transcription

OpenAI covers text-to-speech from the same account as everything else you build on it. The tts-1 and tts-1-hd endpoints are the simple character-billed options at $15 and $30 per million characters. gpt-4o-mini-tts is the newer token-billed model with streaming output, working out to roughly $0.015 per minute of generated audio. The standout is GPT-Realtime-2, OpenAI’s first voice model with GPT-5-class reasoning, which handles speech-to-speech directly — you send audio and reasoning and speech come back together, removing the separate text-to-speech step for agent builds that want one model.

Why it wins: The simplest path if you already build on OpenAI, and the cleanest on-ramp to speech-to-speech agents where reasoning and voice live in one model.

Limitations: The Realtime models lock you to OpenAI’s stack, and dedicated agent APIs such as Cartesia still lead on pure time-to-first-audio. Voice cloning is not offered on these endpoints.

Best for: Products already in the OpenAI ecosystem, and teams building speech-to-speech agents rather than a best-of-breed pipeline.


7. Inworld — best steerable real-time voice with one billing account

Type: Real-time Flagship: Realtime TTS-2 Pricing: From roughly $5 per million characters at scale (Inworld) Key features: Sub-200ms latency, REST and WebSocket streaming, 100+ languages, natural-language steering across emotion, articulation, intonation, volume, pitch, range, speed and vocal style

Inworld Realtime TTS-2, launched in May 2026, tops the Artificial Analysis real-time arena and pairs low latency with an unusual amount of control: you steer delivery across eight dimensions using plain-language instructions rather than choosing a fixed voice. Inworld bundles TTS, speech-to-text, a Realtime API for voice-to-voice interaction and LLM routing across 220+ models behind one API and one bill, which suits teams that want the whole voice stack from a single vendor. Its roots in game NPC dialogue make it a strong pick for interactive and gaming use.

Why it wins: Real-time latency plus fine-grained, natural-language voice steering and a single-vendor voice stack, at aggressive per-character pricing at scale.

Limitations: The all-in-one platform is most valuable if you use several of its pieces; as a pure TTS endpoint it competes on a crowded field. Bundled routing adds a layer of abstraction over the underlying models.

Best for: Interactive and gaming voice, and teams that want steerable real-time speech plus the rest of the voice pipeline from one account.


8. Hume — best for emotion-critical delivery

Type: Expressive Flagship: Octave 2 (text-to-speech) and EVI (empathic voice interface) Pricing: Usage-based API, with plans from $3 to $500 per month; Octave 2 cut costs roughly 50% versus the prior generation Key features: Explicit emotional control via instructions, context-aware delivery, real-time empathic voice interface

Hume splits its offering into Octave, a context-aware text-to-speech model, and EVI, a real-time speech-to-speech model that detects emotional cues and responds in kind. Its differentiator is explicit emotional control: you direct delivery with instructions rather than relying on voice choice alone. The honest trade-off is that this surface is genuinely useful when emotion is the point, but for most workloads ElevenLabs or Google deliver enough expressiveness through voice selection alone, so Hume’s controls are overkill unless emotional fidelity is the requirement.

Why it wins: The most direct programmatic control over emotional delivery, in both text-to-speech and speech-to-speech, for applications where how something is said matters as much as what is said.

Limitations: On raw naturalness it trails the quality leaders, and the emotional-control surface adds complexity you do not need for neutral narration or standard agents.

Best for: Wellness and therapy apps, character voices and empathetic agents where emotional delivery is the load-bearing feature.


9. Microsoft Azure AI Speech — best for enterprise compliance and language breadth

Type: High-volume Flagship: Neural and Neural HD voices Pricing: $15–16 per million characters (neural); $22 per million (Neural HD, cut from $30 in March 2026); $24 per million for custom voice; commitment tiers cut the effective rate to about $7.50 per million; $200 credit plus 500K characters per month free Key features: 500+ voices across 140+ languages, custom neural voice, HIPAA BAA, SOC certification and container deployment

Microsoft Azure AI Speech offers the widest language coverage and the strongest compliance story of the mainstream APIs. Its Neural HD voices approach ElevenLabs long-form quality, and unlike Amazon Polly, Azure supports custom voice cloning at $24 per million characters. For organisations standardised on Microsoft — Entra ID, Azure governance, in-region data — it puts high-quality synthesis behind identity, logging and data-residency controls you already run, with container deployment available for on-premise needs.

Why it wins: The broadest language coverage and the deepest compliance and enterprise-governance posture, inside the Microsoft cloud.

Limitations: Streaming latency is moderate rather than class-leading, and the best value assumes you are already an Azure organisation with the volume to reach a commitment tier.

Best for: Regulated industries, healthcare and multilingual telephony where compliance and language breadth outweigh peak latency.


10. Amazon Polly — best for high-volume AWS-native synthesis

Type: High-volume Flagship: Neural and Generative voices Pricing: $4 per million characters (standard), $16 (neural), $30 (generative), $100 (long-form); 5M standard or 1M neural characters free for 12 months Key features: Mature SDKs, 99.9% uptime SLA, deep AWS integration, LLM-based Generative voices

Amazon Polly is the high-volume workhorse for AWS-native applications. Its Generative voices use LLM-based synthesis at $30 per million characters, though they are English-only with no streaming, while standard neural voices at $16 per million cover the everyday case cheaply. Polly’s advantages are operational: mature SDKs, a 99.9% uptime SLA and cost predictability inside AWS. It has no voice cloning, which rules it out for branded or custom-voice work.

Why it wins: The most predictable, AWS-native path for large-scale synthesis, with a strong SLA and the lowest entry price for standard voices.

Limitations: No voice cloning, English-only Generative voices, and quality that trails the leaderboard leaders. It is an infrastructure choice, not a quality one.

Best for: IVR, accessibility features and any high-volume application where AWS integration and cost predictability beat cutting-edge quality.


Feature comparison: the full matrix

FeatureElevenLabsCartesiaDeepgram Aura-2RimeGoogle TTSOpenAIInworldHumeAzureAmazon Polly
TypeExpressive + real-timeReal-timeReal-timeReal-timeHigh-volumeExpressive + S2SReal-timeExpressiveHigh-volumeHigh-volume
StreamingYes (Flash)YesYesYesYesYesYes (WebSocket)YesYesNeural only
Lowest latency~75ms (Flash)~40–82ms~90mssub-200msModerate~160ms (Realtime)sub-200msModerateModerateNear real-time
Voice cloningYesYes (3s)NoNoYes (custom)NoYesYesYes (custom)No
Emotion / SSML controlAudio tagsYesLimitedProsody-tunedYesLimitedYes (8 dimensions)Yes (explicit)Yes (SSML)SSML
Languages70+40+710+70+ (Gemini)50+100+Multiple140+30+
On-premise optionEnterpriseNoEnterpriseYesNoNoNoNoContainersNo
Free tierCredit-basedYes$200 creditStarter minutes$300 + 1M/moPay-goTrial10K chars/mo$200 + 500K/mo12 months
Trains on your dataNo (default)NoNoNoPaid: noNo (default)NoNoNoNo

Data-retention and training policies change; confirm the current terms in each vendor’s documentation before handling regulated or personal data. Latency figures are vendor-reported time-to-first-audio or time-to-first-byte and vary by region and payload — measure on your own traffic before committing.


Pricing comparison: what you’ll actually pay

Per million characters (list, USD)

APIModelPrice per M charsBilling basisNotes
Amazon PollyStandard$4CharacterCheapest entry; lower quality
InworldRealtime TTS-2~$5 (at scale)CharacterVolume pricing; sub-200ms
Google CloudGemini 3.5 Flash TTS~$6 equivalentToken ($6/M tokens)Cheapest top-tier quality
OpenAItts-1 / gpt-4o-mini-tts~$15Character / tokenmini-tts ≈ $0.015/min
AzureNeural$15–16Character140+ languages
Amazon PollyNeural$16CharacterEveryday AWS default
NeuphonicTTS~$17.60CharacterLow-latency developer option
AzureNeural HD$22CharacterHighest Azure fidelity
DeepgramAura-2$30CharacterFlat, predictable
OpenAItts-1-hd$30CharacterHigher-fidelity batch
Google CloudChirp 3 HD$30CharacterCustom voice, disfluencies
RimeCoda / Arcana / Mist$30–50CharacterPhone-agent prosody
CartesiaSonic 3.5~$35CharacterLatency leader
LMNTTTS~$43.60CharacterDeveloper real-time
Amazon PollyLong-form$100CharacterLong narration
ElevenLabsEleven v3 / FlashPlan-boundCreditCommercial from $5/mo

Cost by audio duration

One hour of speech is roughly 9,000 words, about 54,000 characters. The table below estimates character- or token-billed API cost at that rate; subscription and credit tools bundle a fixed allocation per plan instead.

API1 hour10 hours100 hours
Google Gemini 3.5 Flash TTS~$0.54~$5.40~$54
Amazon Polly Neural~$0.86~$8.64~$86
OpenAI tts-1~$0.81~$8.10~$81
Azure Neural~$0.81~$8.10~$81
Deepgram Aura-2~$1.62~$16.20~$162
Cartesia Sonic~$1.89~$18.90~$189
OpenAI tts-1-hd~$1.62~$16.20~$162

Cost strategy: route by job, not loyalty. Use a premium streaming model such as Cartesia Sonic or Deepgram Aura-2 only for the live conversational turn, push high-volume, latency-tolerant synthesis (notifications, batch narration) to a cheap cloud model like Google Gemini Flash TTS or Amazon Polly, and self-host an open-weight model for privacy-sensitive or offline work. The blended rate, not the headline number, is what moves your bill at scale — and on token-billed models, estimate from audio duration rather than text length.


Latency: streaming, time-to-first-audio and voice agents

For anything a human waits on in real time, the number that matters is time-to-first-audio: how long after you send text the first sound comes back. Below roughly 300ms, turn-taking in a conversation feels natural; above it, the agent feels laggy.

APITime-to-first-audio / byteBest real-time use
Cartesia Sonic Turbo~40msPhone agents, live turn-taking
ElevenLabs Flash v2.5~75msConversational AI with top quality
Cartesia Sonic 3.5~82msReal-time, bring-your-own STT + LLM
Deepgram Aura-2~90ms, sub-200ms TTFBVoice agents on the Deepgram stack
OpenAI GPT-Realtime-2~160msSpeech-to-speech reasoning agents
Inworld Realtime TTS-2sub-200msSteerable interactive voice
Rime (cloud / on-prem)sub-200ms / sub-100msEnterprise phone agents
Google / Azure / PollyNear real-timeHigh-volume IVR and telephony

Use streaming, low-latency APIs for voice assistants, phone agents, live translation and gaming dialogue. Use batch, file-return synthesis for notifications, e-learning, audiobooks and marketing audio, where a few seconds of generation time is irrelevant and cost matters more. Independent voice-AI evaluators caution that vendor latency and quality numbers are measured under favourable conditions, so test on your own audio and traffic before paying a premium (Coval).


Use-case recommendations

For real-time voice agents

Winner: Cartesia Sonic 3.5, or Deepgram Aura-2 on the Deepgram stack

Cartesia’s ~82ms end-to-end latency (and ~40ms on Sonic Turbo) leads for text-to-speech-only agents where you supply your own speech-to-text and language model. If you already transcribe with Deepgram, Aura-2 keeps the whole loop on one vendor at a flat $30 per million characters. Alternative: Rime when conversational phone prosody or on-premise deployment is the priority.

For the best voice quality

Winner: ElevenLabs API with Eleven v3

Top-tier naturalness, 70+ languages and the deepest cloning, for produced audio where quality drives revenue. Use Flash v2.5 for the real-time tier of the same account. Alternative: Google Gemini Flash TTS matches premium quality for a fraction of the cost when budget outranks the last few percent.

For the cheapest synthesis at scale

Winner: Google Cloud Gemini 3.5 Flash TTS (~$6/M tokens)

Roughly $0.09 for a 10-minute narration at top-tier quality. Note that Gemini TTS is excluded from Google’s recurring free allowance, so budget the $300 new-account credit for early testing. Alternative: Amazon Polly standard voices at $4 per million characters for AWS-native, quality-tolerant volume.

For emotion-critical delivery

Winner: Hume Octave 2

Explicit, instruction-driven emotional control for wellness apps, characters and empathetic agents where delivery matters as much as content.

For enterprise compliance and language breadth

Winner: Microsoft Azure AI Speech

140+ languages, custom neural voice, HIPAA BAA and SOC certification inside the Microsoft cloud — the safe default for regulated, multilingual telephony.

For speech-to-speech agents

Winner: OpenAI GPT-Realtime-2

When you want reasoning and voice in one model rather than a separate TTS step, GPT-Realtime-2 is the simplest path. For a fully bundled, no-code phone agent, xAI Grok Voice ships one at $0.05 per minute all-in, though its performance claims are self-reported.

For self-hosting and data control

Winner: Kokoro-82M or Chatterbox

Both run locally with no per-character cost. Kokoro-82M (Apache 2.0, 82M parameters) is the best default for speed and edge deployment and runs faster than real time on modest hardware; Chatterbox (MIT) clones from a 10-second clip but is English-only and watermarks output. For permissive cloning, Qwen3-TTS (Apache 2.0) clones from 3-second samples.


How to choose a text-to-speech API

Nine dimensions decide the fit. Rank them for your build before you compare vendors.

Latency — for anything real-time, time-to-first-audio is the load-bearing number; a dedicated agent API like Cartesia or Deepgram changes what feels possible. Streaming protocol — confirm the API supports chunked or WebSocket streaming, not just file return, if a human is waiting on the audio. Price model — character billing ($4–100 per million) is easy to forecast, while token billing on the newest Google and OpenAI models needs estimating from audio duration. Voice library and cloning — do you need a stock catalogue, a custom brand voice, or instant cloning, and does the vendor support it. Languages — coverage ranges from 7 (Deepgram Aura-2) to 140+ (Azure); match it to your markets. Control — SSML, emotion tags and natural-language steering vary widely; Hume and Inworld lead on explicit control. Concurrency and rate limits — high-volume and real-time workloads hit different limits, so check the concurrent-stream caps before launch. Data and compliance — confirm training, retention, HIPAA, SOC and residency terms in writing for regulated data. Self-host option — if audio cannot leave your network, an open-weight model (Kokoro, Chatterbox) or an on-premise vendor (Rime, Azure containers) is the only path.

Our picks are ranked by suitability for a new voice build, weighting latency, quality, streaming, price and language coverage. We anchor every recommendation in vendor-documented figures and independent leaderboards such as the Artificial Analysis Text-to-Speech arena, and we treat vendor latency and naturalness claims as a ceiling to verify against your own traffic, not a floor.


Building a voice agent: where the TTS API fits

A custom voice agent is a three-stage pipeline, and text-to-speech is the last stage. Speech-to-text transcribes what the user said, a language model decides what to say back, and a text-to-speech API speaks the reply. Each stage is a separate API decision, and the latency of all three stacks up into the total response time the user feels.

For the speech-to-text stage, see best AI for transcription, which compares Deepgram, AssemblyAI, Mistral Voxtral and the other speech-to-text APIs on word error rate and streaming latency. For the language-model stage, see best LLM APIs, which ranks OpenAI, Anthropic, Google and the aggregators on price, speed and model quality. To orchestrate the three stages, an agent framework helps — see best AI agent frameworks.

The alternative is to collapse all three stages into one speech-to-speech model such as OpenAI GPT-Realtime-2 or xAI Grok Voice, which removes the separate text-to-speech step entirely. The three-stage pipeline wins when you want to control the exact voice, mix best-of-breed components or cut cost on high-volume synthesis; the single-model approach wins when you want the simplest build and lowest integration overhead.


For choosing a voice tool with a web interface rather than an API — produced voice-overs, a voice library, a reader — see best AI voice generator. For turning speech into text, see best AI for transcription. For voice cloning specifically and the legal landscape around it, see best AI voice clone. For the wider developer stack, see best LLM APIs, best AI agent frameworks and the best AI apps overview.


Frequently asked questions

What is the best text-to-speech API in 2026?

There is no single best text to speech API — the right choice depends on the job. For real-time voice agents, Cartesia Sonic 3.5 leads on latency at roughly 82ms, and Deepgram Aura-2 ($30 per million characters) is the pick for teams on the Deepgram stack. For the best voice quality, the ElevenLabs API with Eleven v3 is the strongest. For the cheapest synthesis at scale, Google Cloud Gemini 3.5 Flash TTS costs about $6 per million output tokens. Match the API layer — real-time, expressive or high-volume — to your build before comparing vendors.

What is the cheapest text-to-speech API?

Among high-quality options, Google Cloud Gemini 3.5 Flash TTS is the cheapest at about $6 per million output tokens — roughly $0.09 for a 10-minute narration. Amazon Polly standard voices are $4 per million characters but lower quality, and Inworld runs from around $5 per million characters at scale. For zero per-character cost, the open-weight Kokoro-82M model runs locally on modest hardware. Google’s recurring monthly free allowance — 4M characters on Standard voices, 1M on WaveNet — is the cheapest way to prototype, though the Gemini TTS models are excluded from it.

Which text-to-speech API has the lowest latency?

Cartesia has the lowest latency, reaching roughly 82ms end-to-end on Sonic 3.5 and around 40ms on Sonic Turbo, thanks to a state-space model architecture that keeps latency consistent even at the 99th percentile. Deepgram Aura-2 hits a 90ms optimised figure with sub-200ms time-to-first-byte, ElevenLabs Flash v2.5 delivers around 75ms, and Rime reaches sub-100ms on-premise. For natural conversational turn-taking, aim for a time-to-first-audio under 300ms.

What is the best text-to-speech API for voice agents?

For text-to-speech-only voice agents where you supply your own speech-to-text and language model, Cartesia Sonic 3.5 and Deepgram Aura-2 lead on streaming latency, and Rime specialises in conversational phone-agent prosody with an on-premise option. If you would rather use one model for the whole pipeline, OpenAI GPT-Realtime-2 handles speech-to-speech directly. For voice agents, streaming latency and turn-taking usually matter more than raw naturalness.

What is the difference between a text-to-speech API and a voice generator tool?

A text-to-speech API is an endpoint you call from code to synthesise speech inside a product, billed per character or token and measured on latency, streaming and control. A voice generator tool is an application with a web interface, a voice library and an editing workflow, aimed at producing voice-overs, audiobooks or narration without writing code. Developers building agents, captions or in-app speech want an API; creators producing content want a tool. This guide ranks the APIs; our best AI voice generator guide ranks the tools.

Can I run a text-to-speech API locally for privacy?

Yes, and open-weight models are now genuinely competitive. Kokoro-82M (Apache 2.0, 82M parameters) runs faster than real time on modest hardware or CPU with no per-character cost, Chatterbox (MIT) clones a voice from a 10-second clip, and Qwen3-TTS (Apache 2.0) offers permissive cloning from 3-second samples. Self-hosting keeps all audio on your own machine, which removes the data-residency and retention concerns of cloud APIs. The trade-off is that you run and scale the infrastructure yourself.

Do text-to-speech APIs support streaming?

Every agent-grade text-to-speech API in 2026 supports streaming — chunked HTTP or WebSocket delivery so audio starts playing before the full sentence is synthesised. Cartesia, Deepgram Aura-2, ElevenLabs Flash, Rime and Inworld all stream, which is what makes sub-100ms perceived latency possible. Older, file-return endpoints that synthesise the whole clip before returning it still exist for batch narration but are not suitable for live conversation, where the first audio byte needs to arrive within a few hundred milliseconds.

Synthesising speech from text you have written is legal. Voice cloning is where the law applies: cloning your own voice or a voice you have documented consent for is legal, while cloning someone without consent can violate laws such as the Tennessee ELVIS Act, New York’s Digital Replicas Law and the FCC’s robocall order. As of 2026, active class actions are testing how voice-AI training and cloning are regulated. If cloning is central to your product, read our best AI voice clone guide for the detail, and treat a consent checkbox as insufficient legal protection.

How much does a text-to-speech API cost for a real product?

Character-billed APIs run $4 to $100 per million characters, and one hour of speech is about 54,000 characters. At that rate, 100 hours a month costs roughly $54 on Google Gemini Flash TTS, $86 on Amazon Polly Neural, or $162 on Deepgram Aura-2. Real-time agent traffic is usually short per turn but high in volume, so model your cost from expected minutes of audio, not documents. Routing the live turn to a premium model and batch synthesis to a cheap one keeps the blended rate down.


This guide is updated as text-to-speech APIs launch and pricing changes. Latency figures are vendor-reported time-to-first-audio or time-to-first-byte and vary by region and payload — we cite the source for each and treat vendor numbers as a ceiling to verify on your own traffic. Data-retention and training terms change frequently; confirm the current policy with each vendor before handling regulated data. Pricing and availability current as of 11 August 2026 and subject to change.