creative
Best AI Subtitle and Dubbing Tools
Thirteen AI subtitle and dubbing tools compared on price, languages, subtitle file export and lip-sync — including which plans release the subtitle file and what dubbing costs per minute.
Quick answer: The deciding question in this category is not accuracy, because every tool here runs a Whisper-class model and clusters at 2 to 3 per cent word error rate on clean English — it is whether you can get the subtitle file out, and on which plan. For free SRT export, Happy Scribe is the only tool here that puts SRT download on its free plan, though that plan covers 10 minutes of media only. For paid work where the file matters, Sonix includes SRT and VTT export on every tier including its $10-per-hour pay-as-you-go option, with no subscription required. For short-form animated captions, Submagic at $19 per month for 15 videos is built for the format, though its “export subtitles” feature produces a green-screen video rather than a text file. For captions inside a full editor, Descript at $16 per person per month on annual billing gives unlimited dynamic captions on every paid tier. For free captions with no tool at all, YouTube generates automatic captions in 99 languages at no cost. For voice dubbing, ElevenLabs is the cheapest published rate, but read the model names: Dubbing v1 is $0.33 per minute watermarked and $0.50 clean across 29 languages, while the newer Dubbing v2, which covers 92 languages, is $2.20 per minute. Neither offers lip-sync as part of Dubbing. For lip-synced dubbing, HeyGen at $29 per month covers 175-plus languages and charges 2 credits per minute without lip-sync or 5 with it.
The caveat that governs this whole category: the headline accuracy numbers describe clean, single-speaker English and nothing else. Independent evaluator Coval reports that entity accuracy on alphanumerics and proper nouns “drops to 50–70% in many providers’ production output,” and that a model posting 5 per cent word error rate on monolingual audio “routinely posts 15–20% WER on Spanish-English or Hindi-English code-switched calls” (Coval, 4 June 2026). Every tool on this page needs a human pass before the captions are published.
Every price on this page is US list price in USD, checked on 24 September 2026. Where a vendor publishes no list price, or where its pricing page could not be read, this page says “data not available” rather than estimating. Vendor accuracy claims are labelled as vendor claims throughout.
This page covers the media-output layer: captions burned into video, subtitle files you can ship, and translated voice dubbing. That is a different job from the two adjacent pages on this site, and the three do not overlap. If you want to turn speech into a transcript — word error rates, speech-to-text APIs, meeting notetakers — that is best AI for transcription. If you want to turn text into another language — documents, glossaries, localisation pipelines, translation quality benchmarks — that is best AI for translation. This page starts where those two finish: you already have speech, and you need timed captions on screen, a compliant subtitle file, or a translated voice track. For the editors these tools sit inside, see best AI video editor; for the voice models underneath the dubbing, see best AI voice generator and best AI voice clone.
The current state of AI subtitles and dubbing: September 2026
Caption accuracy stopped being the differentiator. What separates these tools now is commercial: which plan releases the subtitle file, how translation is metered, and whether dubbing is charged per source minute or per target language.
Five shifts define the market right now.
1. The subtitle file is the paywall. Generating captions is free or near-free almost everywhere. Downloading them as a file is not. VEED states plainly that “Subtitle file download is available on Pro and Enterprise plans” (VEED support). Kapwing’s help centre is equally direct: “Creating and downloading .SRT files is a paid feature, so you will need to upgrade your workspace” (Kapwing). Riverside restricts transcript and SRT download to paid plans (Riverside support). Two tools break the pattern: Happy Scribe includes SRT on its free plan, and Sonix includes SRT and VTT on every plan including pay-as-you-go. If your workflow ends in a file rather than a rendered video, that single line in the pricing table matters more than any accuracy claim on the page.
2. Word error rate plateaued, and the remaining errors are the expensive ones. Coval’s June 2026 review puts the leading speech-to-text providers within 1 to 2 percentage points of each other on LibriSpeech and FLEURS, at roughly 2 to 3 per cent on clean English. Microsoft MAI-Transcribe-1 posts 3.8 per cent average word error rate across 25 languages on FLEURS; AssemblyAI Universal-3 Pro posts a 5.6 per cent mean, and ranks third on the Artificial Analysis AA-WER v2.0 AgentTalk subset at 2.3 per cent; NVIDIA Parakeet-TDT-0.6B-v3 is the strongest open-source entry at 6.34 per cent average on the HuggingFace Open ASR Leaderboard (Coval). The failures that survive are proper nouns, order numbers, drug names and code-switching — precisely the words a viewer notices when a caption gets them wrong.
3. Dubbing is now priced per language, not per video. ElevenLabs documents the model explicitly: “Dubbing is charged per minute of source media, for each language you dub into” (ElevenLabs docs). A ten-minute video into five languages is billed as fifty minutes. Rask AI goes further and bills lip-sync as a second pass: “to translate and lip-sync a 1-minute video into 1 language, you need 2 minutes” (Rask AI help). HeyGen charges 2 credits per minute for translation without lip-sync, 5 with it, and 10 in Precision mode (HeyGen help centre, cited, not read). Budget on target languages multiplied by source minutes, not on video count.
4. ElevenLabs rebuilt its dubbing model in May, and still does not do lip-sync. ElevenLabs published Dubbing v2 on 28 May 2026, describing a model that conditions on the original performance rather than a transcript, with sync-aware translation and one-click YouTube localisation across 90-plus languages (ElevenLabs). Its documentation is unambiguous about the gap: “ElevenLabs does not offer lip syncing as part of Dubbing”, though lip-sync “is available in Image & Video, Flows, and Studio via third party models” (ElevenLabs docs). Anyone comparing dubbing tools on voice quality alone will pick ElevenLabs; anyone who needs mouths to match needs HeyGen, Synthesia, Rask AI or Deepdub instead.
5. The platforms now dub for free, and the performers are objecting. YouTube auto-dubbing covers 34 source languages, with English dubbing into 19 targets, and is “enabled by default for eligible creators” (YouTube Help). Its expressive speech variant, which replicates the original pitch and intonation, covers 14 languages, and a lip-sync pilot is in early access with selected channels (same source). On the other side of the trade, Amazon’s Prime Video began an AI dubbing pilot on 5 March 2025 across 12 licensed titles in English and Latin American Spanish, framed as titles that “would not have been dubbed otherwise” and paired with human localisation professionals for quality control (Amazon). When AI beta dubs appeared on several anime titles in late November 2025 they were withdrawn within about a week after voice-actor and fan objection — a sequence documented in trade and fan press rather than by any company statement, and worth treating as such (Anime News Network, cited, not read). The United Voice Artists coalition of more than 20 guilds across Europe, the Americas and beyond campaigns against synthetic voice replacement (cited, not read).
What the accuracy evidence actually shows
The 99 per cent standard is a vendor number, not a regulation
Caption vendors routinely cite “99 per cent accuracy” as an industry standard attributed to the Described and Captioned Media Program. That attribution is wrong. The DCMP Captioning Key states the quality goal as “Errorless captions are the goal for each production,” alongside consistency, clarity, readability and equality, and describes itself as “consistent with the 2014 mandates by the Federal Communications Commission” (DCMP). No percentage appears. The 99-per-cent-or-15-errors-per-1,500-words formulation originates with caption vendors and has been repeated until it reads like a standard. The genuine regulatory benchmark is the FCC’s four caption quality standards, covered further down this page.
Vendor accuracy claims, and one vendor contradicting itself
| Tool | Vendor accuracy claim | Note |
|---|---|---|
| VEED | Up to 99.9 per cent | Marketing claim, no methodology published |
| Kapwing | 99 per cent | Marketing claim |
| Maestra | 99 per cent | Marketing claim |
| Riverside | 99 per cent | Marketing claim |
| Sonix | 99 per cent | Sonix’s own guide elsewhere puts real-world accuracy at 85 to 98 per cent |
| Rev | 95 to 96 per cent AI, 99 per cent human | Two-tier claim, human service priced separately |
| Happy Scribe | Up to 85 per cent AI, 99 per cent human | The most conservative AI claim of any tool here |
| Submagic | Data not available | No numeric claim found on a vendor page |
| CapCut | Data not available | No numeric claim found |
Happy Scribe is the outlier worth noticing: it claims up to 85 per cent for its AI subtitle generator where competitors running comparable models claim 99 (Happy Scribe, cited, not read). Sonix is the other: its pricing page claims 99 per cent while its own cost guide concedes 85 to 98 per cent in real use (Sonix; Sonix guide, cited, not read). Treat every figure in that table as marketing, and budget for a review pass regardless of which tool you pick.
The best AI subtitle generators compared
1. Happy Scribe — best free SRT export, and the most conservative accuracy claim
Happy Scribe is the only tool in this comparison that lets you download an SRT file without paying. Its free plan includes a 10-minute trial of AI transcription, subtitling and translation, with exports in DOCX, TXT, SRT and watermarked MP4 (Happy Scribe help centre, page updated 23 July 2026). Paid plans run Basic at $17 per month or $8.50 per month on annual billing for 120 minutes, Pro at $29 or $19 annual for 600 minutes and 3 seats, and Business at $89 or $59 annual for 6,000 minutes and 5 seats. Extra minutes are $0.20 and extra seats $29 per month. AI subtitling covers 150-plus languages, human subtitling 60-plus (cited, not read). The wider export set — VTT, STL, XML, FCPXML, EDL — starts at Pro.
Limitation: credits expire. The help centre states that “Credits refresh every billing period and don’t carry over if unused,” which punishes irregular production schedules.
2. Sonix — best for buying subtitle files without a subscription
Sonix is the only tool here with a genuine pay-as-you-go path to a subtitle file: $10 per hour, no monthly fee, with SRT and VTT export included (Sonix). Subscriptions run Core at $25 per month or $275 per year for 5 hours, Advanced at $50 or $550 for 20 hours and Pro at $80 or $880 for 40 hours. It covers 54-plus languages for transcription and translation. Burned-in subtitles are “included at no extra cost on subscription plans, but will still be charged separately on the pay-as-you-go plan” (Sonix help centre). A 30-minute free trial is available; there is no permanent free plan. Overage is $10 per hour on every plan.
Limitation: included hours belong to the workspace, not the seat. Extra seats cost $25 per month and add no additional hours, so a team of four on Core shares the same five hours one person would get.
3. Descript — best captions inside a real editor
Descript puts unlimited dynamic captions on every paid tier and edits video by editing the transcript. Plans are Hobbyist at $16 per person per month on annual billing or $24 monthly with 10 media hours and 400 AI credits, Creator at $24 or $35 with 30 media hours and 800 credits, and Business at $50 or $65 with 40 media hours and 1,500 credits (Descript). Transcription covers 25 languages. Caption translation reaches 61 languages, audio dubbing 30, and native-sounding AI speakers only 14 (Descript help centre, cited, not read).
Limitation: translation proofread and native-sounding speakers are Business-tier only, so the two features that make translated output publishable sit behind the $50 plan.
4. Submagic — best for short-form animated captions
Submagic is built narrowly for vertical short-form, with styled animated captions as the product rather than a feature. Plans are Starter $19 per month for 15 videos, Professional $39 for 40 and Business $69 for 100, dropping to roughly $12, $23 and $41 per month equivalent on annual billing; the free tier gives 3 watermarked videos per month (cited, not read — Submagic). Its help centre verifies 48-plus languages, in an article last updated in 2023 (Submagic); a newer 100-plus marketing claim could not be verified on a vendor page.
Limitation: Submagic’s “Export Only Subtitles” produces captions rendered onto a green or blue screen video for compositing, not an SRT text file (Submagic). Whether a plain SRT download exists at all, and on which plan, is data not available. Export resolution is also plan-gated at 1080p, 2K and 4K respectively.
5. Kapwing — best browser editor for teams, if you pay
Kapwing generates captions across 155 language variants and translates subtitles into 100-plus (cited, not read). Pro is $16 per member per month billed annually, or $24 billed monthly; the Business price is data not available. The free tier is tightly bounded: uploads capped at 250MB, exports capped at 4 minutes with a watermark at 720p, 3 projects per folder, and projects deleted after 3 days (Kapwing).
Limitation: SRT download is paid-only, and free-workspace projects are auto-deleted after 3 days with each new project locking the oldest.
6. VEED — the most polished editor, with the file behind a paywall
VEED is the most complete browser-based subtitle editor here, claiming 125-plus languages and up to 99.9 per cent accuracy (cited, not read). Its Creator plan “starts at $20 per month” (cited, not read); the remaining tier prices are data not available, because veed.io/pricing is JavaScript-rendered and returned no readable prices on 24 September 2026. Billing is per seat, so every collaborator adds a charge (VEED support).
Limitation: beyond the Pro-only subtitle download, VEED cannot apply two different subtitle styles within one project — the documented workaround is to export the video and re-import it (VEED support).
7. Rev — best when captions must be right rather than cheap
Rev is the only tool here that sells a human pass as a product line rather than an upsell. AI services start at $0.25 per minute. Human English captions are $1.99 per video minute, Spanish captions $3.25, and global subtitles run $6.49 to $15.99 per video minute by language, with Korean and Japanese at the top of that range. Add-ons: Rush at $1.25 per minute, Premium Captions at $1.50, and burned-in captions at $0.30 (Rev support). The free plan includes up to 45 minutes of AI transcription per month. Subscriptions are Essentials at $29.99 and Pro at $59.99 per user per month (cited, not read). AI captions cover 37 languages.
Limitation: per-minute human pricing scales badly. A single hour of video captioned into Japanese costs $959.40 at the $15.99 rate before any add-on.
8. Maestra — best format coverage for broadcast handoff
Maestra exports the widest set of subtitle formats here — SRT, VTT, SCC, STL, CAP, TXT, TTML and SBV — across 125-plus languages, which matters if you are delivering to a broadcaster rather than a social platform (cited, not read). Pricing is $12 per 60 credits pay-as-you-go, Basic at $39 per month for 360 subtitle minutes, and Premium at $79 for 900 (cited, not read — Maestra).
Limitation: translated subtitles consume roughly double the minutes of same-language subtitles — 180 against 360 on Basic — so a bilingual workflow halves your allowance.
9. Riverside — best if you are already recording there
Riverside bundles captions into a recording platform. Standard is $19 per month or $15 on annual billing at $180 per year; Pro is $29 or $24 at $288 per year; Live is $39 or $34 (Riverside). The free tier gives 2 hours of multi-track recording at 720p with a watermark on all output. AI transcription across 100-plus languages is listed as a Pro feature.
Limitation, and one to verify before buying: the help centre lists SRT download as available on “Pro, Grow, Webinar, and Business” — legacy plan names that no longer match the current Standard/Pro/Live line-up, so whether Standard includes SRT is unresolved (Riverside support). Only TXT and SRT are offered; there is no VTT. Transcription is also not generated for screen-share audio, Media Board files or footage uploaded directly to the editor.
The tools we could not price
Three widely recommended tools publish no verifiable USD list price, which is itself a finding. CapCut states that pricing “varies by region and platform” and directs users to the in-app subscription page, so no list price exists to quote (CapCut); its auto-captions are free to generate and export on the web, and it does export standalone SRT. Opus Clip publishes plan names — Free, Starter, Pro, with Pro at 300 processing minutes per month — but its prices could not be read; the free tier gives 60 minutes with a watermark. Zubtitle offers 2 watermarked videos per month free and a Standard plan with 10 credits plus up to 20 rollover credits, at a price that is data not available; note that one video consumes one credit regardless of length, with a 30-minute cap.
The best AI dubbing tools compared
1. ElevenLabs — cheapest on the old model, dearest on the new one, and no lip-sync
ElevenLabs is the default for translated audio where the picture does not need to match. Pricing splits by model: Dubbing v1 is $0.33 per minute with a watermark and $0.50 without, across 29 languages, while Dubbing v2 — the May 2026 rebuild, covering 92 languages — is $2.20 per minute. Dubbing Studio is $0.50 per minute (ElevenLabs). Subscriptions run Free at $0 with 10,000 credits, Starter $6 with 30,000 — the first tier including Dubbing Studio — Creator $22 with 121,000, Pro $99 with 600,000, Scale $299 with 1.8 million and 3 seats, and Business $990 with 6 million and 10 seats (ElevenLabs). ElevenLabs’ documentation puts Dubbing v2 at more than 100 languages and Dubbing v1 at more than 90, with a table of about 130 entries counting regional variants (ElevenLabs docs); its pricing page states 92 and 29 for the two models. Instant voice cloning starts at Starter, professional cloning at Creator.
Limitation: no lip-sync within Dubbing, stated plainly in its own documentation, and the headline $0.33 rate buys the older 29-language v1 model rather than the 92-language v2 at $2.20. Every target language is also billed separately per minute of source.
2. HeyGen — best lip-synced dubbing for creators
HeyGen covers 175 languages and dialects with lip-sync and voice cloning. Plans are Free at 3 videos per month up to 1 minute, Creator $29 per month or $24 annual with 600 credits and videos up to 30 minutes, Pro $49 with 1,000 credits, and Business $149 plus $20 per seat with 1,500 credits and videos up to 60 minutes (HeyGen). Pro tiers scale from 1,000 credits at $49 to 100,000 at $4,300. Translation costs 2 credits per minute without lip-sync, 5 with, and 10 in Precision mode (cited, not read). Monthly subscribers’ unused credits roll over one further month; annual subscribers’ accumulate to renewal.
Limitation: the translation script can only be edited and proofread from the Pro tier upward, so Creator users ship whatever the model produced.
3. Rask AI — broad language coverage, long render times
Rask AI covers 135-plus languages for dubbing with voice cloning across 32. Plans are roughly Creator $60 per month for 25 minutes, Creator Pro $150 for 100, and Business $750 for 500, with extra minutes at $3 (cited, not read). Lip-sync sits on Creator Pro and above and is billed as a separate pass.
Limitation: Rask publishes its own lip-sync render times and they are severe — a 4-minute 1080p video takes about 58 minutes, a 10-minute 1080p video about 4 hours, a 10-minute 4K video about 16 hours, and a 10-minute 1080p video with three speakers about 10 hours 20 minutes (Rask AI help).
4. Synthesia — best for corporate and training video
Synthesia sells dubbing as part of an enterprise video platform with lip-sync included. Basic is $0 with 1,200 credits a month, quoted as 10 minutes of AI dubbing; Starter $29 per month with 1,200 credits and 48 minutes of dubbing; Creator $89 with 3,600 credits and 140 minutes; Enterprise is custom, and the pricing page advertises plans from $18 per month on annual billing (Synthesia). The video translator claims 140-plus languages (cited, not read), while the Enterprise tier advertises one-click translation into 80-plus.
Limitation: credits are a single shared pool across video generation, dubbing and assets, so dubbing competes with everything else you make. Custom avatar creation via Studio Express-1 is a separate $1,000 per year and annual-plan-only, and can take up to 10 days to process.
5. Dubverse — cheapest transparent per-minute dubbing
Dubverse prices in credits: Pro at $0.36 per credit monthly or $0.18 half-yearly, and Supreme at $0.60 or $0.30, both starting at 50 credits with no card required. Dubbing burns 4 credits per minute, text-to-speech 2, and subtitles 1 (cited, not read), which works out to roughly $1.44 per dubbed minute on Pro monthly and about $0.72 on half-yearly billing. It covers 60-plus languages with 500-plus speakers.
Limitation: lip-sync is Enterprise-only and voice cloning starts at Supreme, so the cheap tier is voice-over dubbing and nothing more. Note also that the price list read for this page was last modified in 2024 and may lag the live pricing page.
6. Vozo — lip-sync metered separately
Vozo covers 160-plus languages, with Translate and Dub spanning 111 source and 165 target languages. Plans are Free with 20 AI points, Creator $29 per month with 150 points or roughly 50 dubbing minutes, and Studio $99 with 600 points, roughly 200 dubbing minutes, 60 lip-sync minutes and 3 seats (cited, not read).
Limitation: Vozo’s own documentation warns that “your projects will be automatically deleted 90 days after your subscription expires,” and that the legacy web Pro plan “supports subtitle translation only and does not include AI Points,” excluding dubbing and lip-sync entirely (Vozo docs).
7. Camb.ai — best for long-form under one subscription
Camb.ai covers 150-plus languages with its MARS cloning model, which needs 2 to 3 seconds of reference audio. Tiers run $0, $5, $20, $75, $250 and $900, with video length capped at 5 minutes on free, 20 minutes on the $5 to $75 tiers, and unlimited from $250 (cited, not read). Cloned voice allowances scale from 1 to unlimited across the same tiers.
Limitation: the hard duration caps below $250 rule it out for anything longer than a short segment, and its “lip sync alignment” describes timing matched to mouth movement rather than visual lip re-rendering.
8. The enterprise tier: Deepdub, Papercup and Panjaya
Three vendors serve broadcast and media localisation and publish no self-serve pricing at all. Deepdub offers full lip-sync and voice cloning across 100-plus languages with a 14-day trial of roughly 10 minutes, then time-based packages on quote (Deepdub). Papercup, now the AI dubbing layer inside RWS, states that cost “depends on the required turnaround times, complexity of the content, and the type of technology used” and quotes per project; its exact language count is data not available (Papercup). Panjaya, known for lip-sync-first dubbing, publishes nothing fetchable — pricing, language count and lip-sync terms are all data not available. If you are dubbing a catalogue rather than a channel, these are the vendors to brief, and you will be negotiating rather than subscribing.
Free and built-in options
Before buying anything, check whether the platform you publish on already does this.
| Platform | Captions | Dubbing | Cost | Languages |
|---|---|---|---|---|
| YouTube | Automatic, auto-published | Auto-dubbing, on by default for eligible creators | Free | 99 caption languages; 34 source languages for dubbing |
| Adobe Premiere Pro | Speech to Text, included in subscription | No | No extra charge | 18 transcription languages; caption translation into 27 |
| Apple Final Cut Pro | Transcribe to Captions | No | Included with app | US English only, Apple silicon only |
| Zoom | Automated live captions | No | On by default for paid accounts | 49 languages and dialects |
| Microsoft Teams | Same-language live captions | No | Included; translated captions need Premium or Copilot | Data not available |
| Vimeo | Automatic captions from Standard tier up | No | Caption file upload free on all plans | 99 claimed |
| Canva | Auto-generated editable captions | No | Pro feature | Data not available |
YouTube is the single biggest free option and the most misunderstood. Automatic captions cover 99 languages from Afrikaans to Zulu, and Google is explicit that “quality of the captions may vary” and that speech may be misrepresented “due to mispronunciations, accents, dialects, or background noise,” with creators advised to “always review” (YouTube Help). Expressive captions, which mark intensity and non-speech sounds, are English-only. Live-stream automatic captions are English-only, normal-latency only, and rolling out to channels above 1,000 subscribers.
YouTube auto-dubbing covers 34 source languages, with English dubbing into 19 targets including Arabic, Bengali, Hindi, Japanese, Korean, Portuguese, Spanish and Ukrainian; most non-English sources dub only into English (YouTube Help). Videos over 120 minutes are ineligible, as are videos with little speech, undetectable source language or speech judged too fast. Dubs cannot be edited — which is the reason to pay for a tool if the wording matters.
Adobe Premiere Pro is the strongest free-if-you-already-subscribe option. Adobe states Speech to Text is “included in your Premiere subscription. There is currently no set limit for fair and reasonable use by individual subscribers,” across 18 transcription languages with caption translation into 27, running offline from version 22.2 onward and exporting SCC, MCC, STL, DFXMP or burn-in (Adobe).
Final Cut Pro is more limited than its reputation suggests. Apple’s documentation states that generating captions “requires a Mac with Apple silicon and is available in U.S. English only” (Apple). Reports of expansion into other languages are not reflected in current Apple documentation.
Open-source and self-hosted
If your audio cannot leave your machine, or your volume makes per-minute pricing absurd, the open-source stack is genuinely competitive — the same Whisper family sits under many of the paid tools above.
| Project | Licence | What it gives you | Hardware |
|---|---|---|---|
| OpenAI Whisper | MIT | Six model sizes, reference implementation | tiny 39M at about 1GB VRAM to large 1550M at about 10GB |
| faster-whisper | MIT | Up to 4x faster at the same accuracy, lower memory | CUDA 12 and cuDNN 9 for GPU; runs on CPU with int8 |
| WhisperX | BSD-2-Clause | Word-level timestamps and speaker diarization | Under 8GB GPU memory for large-v2 |
| Subtitle Edit | GPL-3.0 | Desktop subtitle editor driving several Whisper engines | Windows-native, Avalonia build for Linux |
Whisper’s model sizes trade speed against accuracy predictably: tiny at 39M parameters needs about 1GB of VRAM and runs roughly 10 times real time, large at 1550M needs about 10GB and runs at roughly real time, and turbo at 809M needs about 6GB at roughly 8 times real time as an optimised large-v3 with “minimal degradation in accuracy” — but turbo “is not trained for translation” (OpenAI). OpenAI publishes per-language error rates as a chart rather than a headline figure; the “12 per cent WER” number circulating for Whisper does not appear in its repository.
faster-whisper is the practical default. SYSTRAN reports it is “up to 4 times faster than openai/whisper for the same accuracy while using less memory,” with published benchmarks on 13 minutes of audio using large-v2 on an RTX 3070 Ti: 2 minutes 23 seconds at 4,708MB for the reference implementation against 1 minute 3 seconds at 4,525MB for faster-whisper, dropping to 59 seconds at 2,926MB with int8 quantisation and 16 seconds with int8 batching (faster-whisper).
WhisperX is the one to use for subtitles specifically, because subtitle timing is its whole point: 70 times real time with batched large-v2, wav2vec2 forced alignment for word-level timestamps, pyannote diarization for speaker labels, and voice-activity preprocessing that “reduces hallucination & batching with no WER degradation” (WhisperX). It took first place in the Ego4d transcription challenge and was published at INTERSPEECH 2023. Default alignment models cover English, French, German, Spanish and Italian, with others available from HuggingFace.
The trade-off: none of these produce styled animated captions, none dub, and all of them require you to own the pipeline. NVIDIA Parakeet-TDT-0.6B-v3 is currently the strongest open-source model on raw accuracy at 6.34 per cent average word error rate on the HuggingFace Open ASR Leaderboard — respectable, but roughly double the best commercial figures (Coval).
Pricing comparison: what you will actually pay
Subtitle and caption tools
| Tool | Entry price | Best-value paid tier | Subtitle file export | Languages |
|---|---|---|---|---|
| Happy Scribe | Free, includes SRT | Basic $8.50/mo annual, 120 min | SRT on free; VTT, STL, XML from Pro | 150+ |
| Sonix | $10/hour pay-as-you-go | Core $275/year, 5 hrs/mo | SRT and VTT on every plan | 54+ |
| Descript | Free tier | Hobbyist $16/person/mo annual | Included; unlimited dynamic captions | 25 transcription |
| Kapwing | Free, watermarked | Pro $16/member/mo annual | Paid plans only | 155 variants |
| Submagic | Free, 3 videos watermarked | Starter $19/mo, 15 videos | Green-screen video, not SRT | 48+ verified |
| VEED | Free, watermarked | Creator from $20/mo | Pro and Enterprise only | 125+ |
| Rev | Free, 45 AI min/mo | AI from $0.25/min | All formats included with order | 37 AI captions |
| Maestra | $12 per 60 credits | Basic $39/mo, 360 min | SRT, VTT, SCC, STL, CAP, TTML, SBV | 125+ |
| Riverside | Free, watermarked | Standard $15/mo annual | Paid plans, TXT and SRT only | 100+ on Pro |
| CapCut | Free web auto-captions | Data not available | SRT export available | 20 to 30 |
| Opus Clip | Free, 60 min watermarked | Data not available | SRT and VTT claimed | 40+ |
| Zubtitle | Free, 2 videos/mo | Data not available | SRT and TXT | 60+ |
Dubbing tools, by cost per minute where published
| Tool | Cost per dubbed minute | Subscription entry | Lip-sync | Languages |
|---|---|---|---|---|
| ElevenLabs | v1 $0.33 watermarked, $0.50 clean; v2 $2.20 | Starter $6/mo | No | v1 29, v2 92 |
| Dubverse | About $1.44 on Pro monthly | 50 credits, no card | Enterprise only | 60+ |
| Rask AI | About $2.40 at Creator Pro rate; extra minutes $3 | Creator about $60/mo | Yes, billed as a second minute | 135+ |
| HeyGen | 2 credits/min, 5 with lip-sync | Creator $29/mo | Yes | 175 |
| Synthesia | Metered from shared credit pool | Starter $29/mo | Yes | 140+ claimed |
| Vozo | Metered in AI points | Creator $29/mo | Yes, metered separately | 160+ |
| Camb.ai | Not published per minute | $5/mo | Alignment only | 150+ |
| Descript | Metered in AI credits | Hobbyist $16/person/mo | No | 30 dubbing |
| Deepdub | Quote only | Contact sales | Yes, full lip-sync | 100+ |
| Papercup | Quote only | Contact sales | Data not available | Data not available |
| YouTube | Free | Free | Experimental pilot | 34 source |
Subtitle file formats, and when each one matters
Format choice is the difference between a file that works and a redelivery. The short version:
- SRT is the universal baseline. Plain text, no styling, accepted by every social platform and every player. If you only ever need one format, it is this one.
- VTT (WebVTT) is the web standard, supports basic positioning and styling, and is what HTML5 video expects. Note that Riverside offers no VTT at all.
- SCC is the broadcast caption standard in the US and carries the positioning and roll-up behaviour broadcasters expect. Premiere Pro and Maestra both export it.
- STL (EBU-STL) is the European broadcast equivalent. Maestra and Premiere Pro export it; most creator tools do not.
- TTML and DFXP are XML-based and used by streaming platforms and Netflix-style delivery specifications.
- Burned-in (open) captions are pixels, not text. They cannot be switched off, cannot be translated afterwards, and are not accessible to screen readers — but they are the only reliable option for silent autoplay on social feeds.
The practical rule: burn in for social, ship SRT or VTT for the web, and ask the recipient for their delivery specification before you render anything for broadcast. If a tool cannot export the format your distributor requires, no amount of caption accuracy rescues it.
Captions and the law: what compliance actually requires
If you are captioning for a university, a public body, a broadcaster or a product sold into the EU, the requirement is not “good captions” — it is a named standard with a date attached. Four instruments matter, and the most commonly cited deadline in this area is now wrong.
ADA Title II, United States — the deadlines moved, and none has passed. The US Department of Justice published its final rule on 24 April 2024, requiring state and local government web content and mobile apps to meet WCAG 2.1 Level AA (ADA.gov). The original compliance dates of April 2026 and April 2027 were extended by an Interim Final Rule published on 20 April 2026. The current dates are 26 April 2027 for public entities with a population of 50,000 or more, and 26 April 2028 for those under 50,000 and for special district governments (ADA.gov; Federal Register). As of 24 September 2026 no ADA Title II web deadline has passed; the first is 26 April 2027. The underlying obligations were not changed, only the dates — and any comparison page still citing an April 2026 deadline as live is describing a rule that was amended five months ago.
European Accessibility Act — already in force. Directive (EU) 2019/882 applies from 28 June 2025 and covers access to audiovisual media services and related consumer equipment (European Commission; application date cited, not read). Its accessibility components — subtitles for deaf and hard-of-hearing users, audio description, spoken subtitles and sign language interpretation — must be fully transmitted in high quality, accurately displayed and synchronised with sound and video, which is a direct constraint on machine-timed captions. Service providers’ facilities lawfully in use by 28 June 2025 have transitional relief until 28 June 2030. The harmonised standard EN 301 549 was updated, announced 7 September 2026 (Accessible EU Centre).
FCC caption quality, United States broadcast — four standards, unchanged. The FCC’s rules at 47 CFR 79.1(j)(2), adopted by Report and Order FCC 14-12 on 20 February 2014, require captions to meet four standards: accuracy, synchronicity, completeness and placement, across prerecorded, live and near-live programming (FCC). Nothing in those four standards changed in 2025 or 2026. Separately, a Third Report and Order on closed captioning display settings carries a compliance date of 17 August 2026 for covered apparatus manufacturers and MVPDs (cited, not read).
WCAG — the success criteria you are actually being measured against. Captions for prerecorded media are 1.2.2, Level A. Captions for live media are 1.2.4, Level AA. Audio description or a media alternative for prerecorded media is 1.2.3, Level A, and audio description alone is 1.2.5, Level AA (W3C). Because ADA Title II mandates WCAG 2.1 AA, all four are in scope for US public entities. The numbering and levels are identical across WCAG 2.0, 2.1 and 2.2.
What this means for tool choice. Automatic captions alone do not meet any of these standards, because none of them tolerate the 2 to 3 per cent error floor on names and numbers that the benchmarks describe. For compliance work, the shortlist narrows to tools that support a reviewed human pass and export a broadcast-grade format: Rev for the human pass at $1.99 per video minute for English, Maestra or Adobe Premiere Pro for SCC and STL delivery, and Happy Scribe for its 99 per cent human subtitling tier. Budget for review time as a line item, not an afterthought.
Best AI subtitle and dubbing tool by use case
Best free subtitle file: Happy Scribe
The only tool here that exports SRT on a free plan, with a 10-minute AI trial and 150-plus languages. Happy Scribe
Best value for regular subtitle work: Sonix
SRT and VTT on every plan including $10-per-hour pay-as-you-go, so you never pay a subscription to retrieve your own file. Sonix
Best for short-form vertical video: Submagic
Purpose-built animated captions at $19 per month for 15 videos — accepting that you get a rendered video rather than a text file. Submagic
Best captions inside an editor: Descript
Unlimited dynamic captions on every paid tier from $16 per person per month on annual billing, with transcript-based editing. Descript
Best for teams collaborating in a browser: Kapwing
Pro at $16 per member per month annually, 155 language variants, strong review workflow — but budget for the paid tier, because the free one deletes projects after 3 days. Kapwing
Best when the captions must be correct: Rev
Human English captions at $1.99 per video minute and a 99 per cent human accuracy claim, which is what compliance work actually requires. Rev
Best broadcast format coverage: Maestra
SRT, VTT, SCC, STL, CAP, TXT, TTML and SBV across 125-plus languages, from $39 per month. Maestra
Best free option overall: YouTube
Automatic captions in 99 languages and auto-dubbing from 34 source languages at no cost, if you accept that you cannot edit the dub. YouTube
Best if you already pay Adobe: Adobe Premiere Pro
Speech to Text included with no set usage limit, 18 transcription languages, caption translation into 27, and SCC, MCC, STL and DFXMP export. Adobe
Cheapest published dubbing rate: ElevenLabs
Dubbing v1 at $0.33 per minute watermarked or $0.50 clean across 29 languages, with Dubbing Studio from the $6 Starter plan. The 92-language Dubbing v2 is $2.20 per minute, and neither includes lip-sync. ElevenLabs
Best lip-synced dubbing: HeyGen
175 languages with lip-sync from $29 per month, at 5 credits per minute with lip-sync against 2 without. HeyGen
Best dubbing for corporate and training video: Synthesia
Lip-sync included, 140-plus claimed languages, and 48 minutes of dubbing on the $29 Starter plan. Synthesia
Cheapest transparent per-minute dubbing: Dubverse
About $1.44 per dubbed minute on Pro monthly billing, or roughly half that on half-yearly, with no card required to start. Dubverse
Best for self-hosted and private workflows: WhisperX
Word-level timestamps and speaker diarization under 8GB of GPU memory, free under a BSD-2-Clause licence, which is the right base for subtitle timing specifically. WhisperX
Best for media localisation at catalogue scale: Deepdub
Full lip-sync and voice cloning across 100-plus languages on quote, alongside Papercup and Panjaya in the enterprise tier. Deepdub
Frequently asked questions
What is the best AI subtitle generator in 2026?
There is no single best tool, because the deciding factor is whether your plan lets you download the subtitle file. Happy Scribe is the best free option because it includes SRT export on its free plan, which no other tool here does. Sonix is the best paid option for file-based work because SRT and VTT are included on every tier including its $10-per-hour pay-as-you-go plan. Submagic is the best choice for animated short-form captions at $19 per month, and Descript is the best choice if you want captions inside a full video editor at $16 per person per month on annual billing.
How much does an AI subtitle generator cost?
Entry pricing runs from free to about $39 per month. Happy Scribe Basic is $8.50 per month on annual billing for 120 minutes, Kapwing Pro and Descript Hobbyist are both $16 per month, Submagic Starter is $19 for 15 videos, VEED Creator starts at about $20, Sonix Core is $25 per month or $275 per year, and Maestra Basic is $39 for 360 subtitle minutes. Pay-as-you-go alternatives are Sonix at $10 per hour, Rev AI services from $0.25 per minute and Maestra at $12 per 60 credits. CapCut, Opus Clip and Zubtitle publish no verifiable US list price.
Can I get an SRT file for free?
Yes, from two routes. Happy Scribe includes SRT download on its free plan alongside DOCX, TXT and watermarked MP4 export, though that free plan covers a 10-minute trial of AI subtitling rather than ongoing use, and VTT is Pro-tier only. Alternatively, run Whisper, faster-whisper or WhisperX locally — all free and open-source — which gives you SRT with no minute limits at all. Most browser tools paywall the file: VEED restricts subtitle download to Pro and Enterprise, Kapwing states SRT creation and download is a paid feature, and Riverside limits transcript and SRT download to paid plans.
How accurate are AI subtitles really?
On clean, single-speaker English, the leading engines sit at roughly 2 to 3 per cent word error rate, and vendors round this to claims of 99 per cent. The accuracy that matters is worse. Independent evaluator Coval reports entity accuracy on alphanumerics and proper nouns dropping to 50 to 70 per cent in production output, and code-switched audio pushing a 5 per cent model to 15 to 20 per cent. Streaming captions typically run 1 to 3 points worse than batch processing. Happy Scribe is the only vendor here that publishes a conservative figure, claiming up to 85 per cent for its AI subtitling.
Is the 99 per cent caption accuracy standard real?
Not as a standard. The 99 per cent figure, often expressed as no more than 15 errors per 1,500 words, originates with caption vendors and is frequently misattributed to the Described and Captioned Media Program. The DCMP Captioning Key states the goal as “Errorless captions are the goal for each production” and contains no percentage. The genuine regulatory benchmark in the United States is the FCC’s four caption quality standards — accuracy, synchronicity, completeness and placement — at 47 CFR 79.1(j)(2).
What is the best AI dubbing tool?
ElevenLabs publishes the cheapest rate, but check which model you are buying: Dubbing v1 is $0.33 per minute watermarked and $0.50 clean across 29 languages, while the 92-language Dubbing v2 is $2.20 per minute. Neither offers lip-sync. If you need the speaker’s mouth to match the new audio, HeyGen covers 175 languages from $29 per month, Synthesia includes lip-sync from its $29 Starter plan, and Rask AI covers 135-plus languages but bills lip-sync as a second minute. For catalogue-scale media localisation, Deepdub, Papercup and Panjaya are enterprise vendors you brief rather than subscribe to.
How much does AI dubbing cost per minute?
ElevenLabs publishes the cheapest rate on its older model, Dubbing v1, at $0.33 per minute watermarked and $0.50 clean; its 92-language Dubbing v2 is $2.20 per minute. Dubverse works out to roughly $1.44 per dubbed minute on Pro monthly billing. Rask AI charges $3 per extra minute above plan allowance. HeyGen and Vozo meter dubbing in credits and points rather than minutes, and Synthesia draws on a shared credit pool. Be aware that most vendors charge per minute of source media for each target language, so a ten-minute video into five languages is billed as fifty minutes.
Does AI dubbing include lip-sync?
It depends on the tool, and the difference is expensive. ElevenLabs and Descript do not offer lip-sync for dubbing at all. HeyGen, Synthesia, Deepdub and Vozo do. Rask AI offers it from Creator Pro and bills it as a separate pass, so a one-minute video translated with lip-sync into one language consumes two minutes of allowance. Dubverse restricts lip-sync to Enterprise. Camb.ai offers “lip sync alignment”, meaning timing matched to mouth movement rather than visual re-rendering of the lips. YouTube’s lip-sync is an experimental pilot limited to selected channels.
Does YouTube dub videos automatically?
Yes, for eligible creators, at no cost. YouTube auto-dubbing supports 34 source languages, and English videos can be dubbed into 19 target languages including Arabic, Bengali, Hindi, Japanese, Korean, Portuguese, Spanish and Ukrainian. Most non-English source languages dub only into English. It is enabled by default for eligible creators, with early access self-enablement in Advanced settings for others. The significant limitation is that dubs cannot be edited, and videos over 120 minutes are ineligible.
Do I need to pay for captions if I only post to social media?
Usually not. CapCut generates and exports auto-captions free on the web, YouTube captions videos automatically in 99 languages, and Descript, Kapwing and VEED all caption on free tiers — the free-tier restrictions are on watermarks, export length and file download rather than on generating the captions. You need to pay when you want an unwatermarked export, a subtitle file you can ship elsewhere, a specific animated style, or translation into another language.
Which subtitle format should I use?
Use SRT unless something specifically requires otherwise: it is plain text and accepted everywhere. Use VTT for HTML5 web video, where basic positioning and styling are supported. Use SCC for United States broadcast delivery and EBU-STL for European broadcast. Use TTML or DFXP for streaming platform delivery specifications. Use burned-in captions for social feeds that autoplay silently, accepting that burned-in text cannot be switched off, translated afterwards or read by assistive technology.
Are automatic captions enough for legal accessibility compliance?
No. Automatic captions alone do not meet WCAG 2.1 Level AA, which ADA Title II mandates for US state and local government web content, because the 2 to 3 per cent error floor on names and numbers fails the accuracy requirement. The relevant success criteria are 1.2.2 Captions Prerecorded at Level A, 1.2.4 Captions Live at Level AA, 1.2.3 Audio Description or Media Alternative at Level A and 1.2.5 Audio Description at Level AA. For compliance work you need a reviewed human pass — Rev at $1.99 per video minute for English or Happy Scribe’s 99 per cent human tier — plus a tool that exports the delivery format your distributor requires.
When is the ADA Title II captioning deadline?
26 April 2027 for public entities with a population of 50,000 or more, and 26 April 2028 for those under 50,000 and for special district governments. These dates replaced the original April 2026 and April 2027 deadlines through an Interim Final Rule published on 20 April 2026. As of 24 September 2026 no deadline has passed. The requirement itself is unchanged: WCAG 2.1 Level AA for web content and mobile applications, which includes captions for prerecorded and live media.
Can I run subtitle generation on my own computer?
Yes, and it is free. OpenAI Whisper is MIT-licensed with six model sizes, from tiny at 39M parameters needing about 1GB of VRAM to large at 1550M needing about 10GB. faster-whisper is up to four times faster at the same accuracy and runs on CPU with int8 quantisation. WhisperX is the best choice for subtitles specifically, because it adds word-level timestamps through forced alignment and speaker diarization while using under 8GB of GPU memory. Subtitle Edit is a free GPL-3.0 desktop editor that drives several of these engines. The trade-off is no animated caption styling, no dubbing and a pipeline you maintain yourself.
Is AI dubbing replacing human voice actors?
It is being deployed alongside them, and contested. Amazon’s Prime Video began an AI dubbing pilot on 5 March 2025 across 12 licensed titles in English and Latin American Spanish, explicitly framed as titles that “would not have been dubbed otherwise” and paired with human localisation professionals for quality control. When AI beta dubs appeared on several anime titles in late November 2025 they were withdrawn within about a week following voice-actor and fan objection. The United Voice Artists coalition of more than 20 guilds campaigns against synthetic voice replacement, and consent, replica control and compensation are the terms under negotiation across the industry.
Conclusion: how to choose in September 2026
Work backwards from the artefact you need to ship.
If you need a subtitle file, the question is which plan releases it. Happy Scribe gives you SRT free; Sonix gives you SRT and VTT on every plan including $10-per-hour pay-as-you-go; Maestra gives you the broadcast formats. Everything else here charges for the download, and VEED, Kapwing and Riverside all put it behind a paid tier.
If you need captions burned into short-form video, Submagic at $19 per month is purpose-built and CapCut is free, and neither will give you a usable text file at the end.
If you need a translated voice track, decide on lip-sync first, because it sets the price. Without it, ElevenLabs is the cheapest published rate — $0.33 to $0.50 a minute on the 29-language Dubbing v1, or $2.20 on the 92-language v2. With it, HeyGen at $29 per month and 5 credits per minute, or Synthesia at $29 with lip-sync included, are the practical options — and budget on source minutes multiplied by target languages, not on video count.
If you need captions that satisfy a legal standard, no automatic tool gets you there alone. Budget for a human pass and check the delivery format before you render.
And if you publish to YouTube, try the free path first: 99 caption languages and 34 source languages of auto-dubbing at no cost, with the single caveat that you cannot edit the dub.
Transcription accuracy has stopped being the deciding factor between these tools; the commercial terms have taken its place. Check the file-export line in the pricing table, check whether dubbing is billed per target language, and check whether your credits expire — those three lines will cost or save you more than any accuracy claim on this page.