Most people who try to dub a video into another language hit two walls. The first is the bill. Cloud dubbing is charged per minute of video and again for every language, so one ten-minute video in three languages costs you thirty minutes. The second is the upload form. If you work under a client NDA or a company policy on external tools, that's where you stop and send an email asking whether you're allowed.
We added dubbing to Demodokos Foundry in version 2.3 because we kept running into both walls ourselves. It runs on your own graphics card. It transcribes, translates and re-voices your video in the original speaker's voice, and none of it leaves your PC.
Key takeaways
- AI dubbing software transcribes the original speech, translates it, and re-voices it in the target language, ideally keeping the original speaker's voice.
- Demodokos Foundry dubs audio and video into 10 spoken languages, with subtitles in 50. The whole pipeline runs locally on a Windows PC.
- Cloud dubbing tools sell a fixed number of dubbed minutes per month, and every language counts separately. Foundry has no minute counter: from $12 a month, you can dub as much as your GPU can process.
- The tradeoffs: you need a capable GPU, Windows 10 or 11, and a 20 to 25 GB model download on first launch. Foundry speaks 10 languages and dubs the audio only, without lip-sync.
- Machine translation still needs a human read-through. Foundry lets you pause the dub and fix single lines, voices and timings before you export.
What does AI dubbing software actually do?
AI dubbing software does four jobs in a row: it transcribes what was said, translates it, generates new speech in the target language, and fits that speech back into the original timing. The quality of a dub depends on all four. A perfect voice with a clumsy translation still sounds wrong, and so does a perfect translation that runs two seconds past the speaker's mouth.
Transcription turns speech into text with timestamps for each word. Foundry uses its own speech recognition model, Foundry ASR v4, which records word-level timing. The next steps depend on that timing.
Translation is where names and jargon go wrong. A literal translation model will happily turn a brand called Apple into "Mela" in Italian. Good dubbing tools let you mark words that must stay as they are.
Voicing is the part people notice first. Older tools used a stock voice, so a French interviewee suddenly sounded like an American radio host. Current tools clone the original speaker from the video's own audio and speak the translation in that voice.
Timing adaptation is the least visible and the hardest part. Translated sentences are often longer or shorter than the original, and German or Spanish versions of an English line frequently run longer. The software has to adjust pacing so each line starts and ends close to where the original did.
Why dub video locally instead of in the cloud?
Local dubbing matters most for two groups: people whose files are not allowed to leave the building, and people who dub often enough that minute limits become a monthly problem.
The first group is bigger than it sounds. It includes unreleased online courses, internal training videos, client interviews, game trailers under embargo, and anything covered by GDPR or a data processing agreement. With a cloud tool, the full video goes to a third-party server before a single word is translated. With Foundry, the video, the transcript, the cloned voice and the finished dub stay on your own drive. You need an internet connection to sign in and verify your license, but not to process anything.
The second group feels it on the invoice. Cloud tools count every minute in every language, and on credit-based plans every regeneration counts too. ElevenLabs charges credits per generation request, not per download, so fixing a line and re-dubbing it draws from the same balance again. A local tool costs the same whether you dub one video this month or forty, and whether you redo a line once or ten times.
There's also no queue. On a strong GPU, Foundry's speech model produces a full hour of finished speech in under five minutes (measured on an RTX 5090).
How many minutes of AI dubbing do you actually get per month?
On cloud dubbing tools, your plan decides how many dubbed minutes you get each month, and every target language counts separately. Most people only notice this once they add a second language.
Here's how the numbers work out. ElevenLabs plans are Starter $6 (30,000 credits), Creator $22 (121,000 credits), Pro $99 (600,000 credits), Scale $299 (1,800,000 credits) and Business $990 (6,000,000 credits). Dubbing costs 3,000 credits per minute for automatic dubbing without a watermark, and 10,000 credits per minute in Dubbing Studio without a watermark. Dubbing Studio is the editor you need to correct a translation. HeyGen's Creator plan costs $29/month, Pro starts at $49/month, and Business costs $149/month plus $20 per seat, with 600 credits on Creator, 1,000 on base Pro and 1,500 on Business. Video translation uses 2 credits per minute for audio dubbing or 5 credits per minute with lip sync.
| Plan | Price per month | Dubbed minutes included per month |
|---|---|---|
| ElevenLabs Creator | $22 | 40 automatic, or 12 in Dubbing Studio |
| ElevenLabs Pro | $99 | 200 automatic, or 60 in Dubbing Studio |
| ElevenLabs Scale | $299 | 600 automatic, or 180 in Dubbing Studio |
| ElevenLabs Business | $990 | 2,000 automatic, or 600 in Dubbing Studio |
| HeyGen Creator | $29 | 300 audio dub, or 120 with lip-sync |
| HeyGen Pro (base) | $49 | 500 audio dub, or 200 with lip-sync |
| HeyGen Business | $149 + seats | 750 audio dub, or 300 with lip-sync |
| Demodokos Foundry Creator | from $12 | No limit |
| Demodokos Foundry Professional | from $39.20 | No limit |
Minutes are counted per target language. Figures are calculated from each provider's published credit rates, September 2026.
These minutes also compete with everything else on the same plan. On ElevenLabs, text-to-speech, speech-to-text, voice cloning and dubbing all draw from the same monthly balance. On HeyGen, credits reset each billing cycle and generally do not roll over, and Creator does not support one-time credit purchases, so the only way to get more credits is to upgrade to Pro. HeyGen's Creator plan also caps a single video at 30 minutes.
How much does AI dubbing cost for a real workload?
The gap between cloud and local dubbing grows with every language and every video you add. Here are three realistic workloads and the cheapest option on each service that covers them. The comparison uses automatic dubbing on ElevenLabs and audio dubbing on HeyGen, because Foundry also dubs the audio without lip-sync.
| Workload | Dubbed minutes per month | ElevenLabs | HeyGen | Demodokos Foundry |
|---|---|---|---|---|
| Weekly YouTube channel: 4 videos × 10 min, 3 languages | 120 | $99 (Pro) | $29 (Creator) | from $12 |
| Twice a week, 5 languages: 8 videos × 10 min | 400 | $299 (Scale) | $49 (Pro) | from $12 |
| 10-hour online course into 4 languages | 2,400 | about $1,200 (API) | about $314 (Business + extra credits) | from $12 |
For the course, ElevenLabs Business covers only 2,000 minutes, so the table uses the API rate of $0.50 per minute for automatic dubbing without a watermark. On HeyGen, Business users can buy extra credits at $0.05 per credit, so the remaining 3,300 credits add $165 to the $149 plan.
Over twelve months, the second workload costs $3,588 on ElevenLabs Scale, $588 on HeyGen Pro, and from $144 on Foundry. And the Foundry price covers music generation, voice generation and the timeline editor too, not just dubbing. If you dub videos for clients, the Professional plan from $39.20 a month is the right license, and it has no minute counter either.
How do you dub a video with Demodokos Foundry?
You dub a video in Foundry by importing it, picking a target language, and letting the pipeline transcribe, translate and re-voice it. You can pause anywhere to review. The full walkthrough is in our 7-minute dubbing tutorial. The short version:
- Import the video or audio file. Foundry accepts both, so a podcast episode works as well as a YouTube upload.
- Choose the target language. You can pick any of the 10 spoken languages. Subtitles can go into any of 50.
- Set hot words. List names, product names and technical terms that should not be translated or should be spelled a specific way. This step prevents most of the embarrassing errors.
- Run the translation and voicing. Foundry transcribes the original with word-level timing, translates it, clones each speaker's voice and adapts the timing of every line.
- Pause and review. Read through the translated lines before or after voicing. You can correct a line, add a pronunciation rule, change a speaking style, swap the voice for a speaker or adjust a single line's timing, then regenerate only that line.
- Export with subtitles. The finished dub exports with styled subtitles in the target language.
Step 5 is where the result goes from passable to publishable. No translation model knows that your recurring character "Chef" is a nickname and not a job title. Two minutes of review catches that, and since nothing is metered, you can regenerate a line as many times as it takes.
How close does the dubbed voice sound to the original?
The dubbed voice is a clone of the original speaker, built from the audio in the video itself. It keeps their timbre and general character in every target language. Foundry's Speech v4 model keeps a voice's identity stable across long recordings, and the same engine drives more than 40 emotions and speaking styles at five intensity levels each. That means you can make a line calmer or more excited without it turning into a different person. The voice cloning and emotion engine post explains how this works.
You can judge it yourself on the Demodokos homepage. It has 7 openly licensed clips, including an anime trailer, a German robotics interview, a French basketball podcast and a Korean unboxing video, with 90 original and dub pairs across 10 languages. All of them were made locally in Foundry.
Clean source audio gives the best results. Loud background music, people talking over each other and noisy field recordings make every step harder, for Foundry and for every other dubbing tool.
What are the limits of dubbing in Foundry?
Foundry's dubbing has clear limits, and it's better to know them before the trial than after it.
- No lip-sync. Foundry replaces the voice track and leaves the picture untouched. For voice-over, tutorials, screen recordings, podcasts, animation and interviews with cutaways, that's all you need. For close-up talking heads where viewers watch the speaker's mouth, a lip-sync tool does something Foundry doesn't.
- Languages: 10 spoken languages: English, German, French, Spanish, Italian, Chinese, Japanese, Russian, Portuguese and Korean. Subtitles support 50.
- Hardware: An NVIDIA GPU with 6 GB of VRAM is a practical minimum, and 12 GB or more is recommended. AMD cards with 8 GB+ work through Vulkan, but AMD support is still experimental. Our GPU guide for local AI audio goes through card choices.
- Platform: Windows 10 and 11 only. There's no Mac or Linux version yet.
- Multiple speakers: Automatic speaker separation for videos with several people talking is part of the Professional plan.
Frequently Asked Questions
Can I try AI dubbing in Foundry for free?
Yes. Demodokos Foundry has a 7-day free trial with full access to the plan you choose. It starts with a $0 PayPal authorization, and you're only charged if you keep the subscription after the trial ends.
Is there a limit on how many minutes I can dub?
No. Foundry has no monthly minute allowance and no credits. You can dub as many videos into as many of the 10 languages as your GPU can process, and regenerate lines as often as you like.
Can I dub a video that is longer than an hour?
Yes. On the Creator plan, you dub long videos in parts of up to 20 minutes and join them on Foundry's timeline. The Professional plan takes files up to 1 hour 30 minutes in one go.
Does Foundry need an internet connection to dub a video?
Foundry needs an internet connection to sign in and verify your license. Transcription, translation, voice cloning and speech generation all run on your own GPU, and your video is never uploaded for processing.
Can I edit the translation before the dubbed voice is generated?
Yes. You can pause the dubbing process to review and correct translated lines, add pronunciation rules and hot words, and change voices or timing for single lines. Only the lines you changed get regenerated.
Can Foundry transcribe audio without dubbing it?
Yes. Transcription is the first step of every dub, and it's also available on its own. The Professional plan adds live dictation from a microphone, system audio or your phone, with speaker separation for recordings that have multiple voices.
Try it on one of your own videos
The quickest way to judge any dubbing tool is to run your own footage through it. Pick a video you know well, dub it into a language you can at least half follow, and listen for the moments where it slips. Start the free trial, then follow the dubbing tutorial. If you're coming from a cloud tool, our ElevenLabs alternatives comparison covers the rest of the voice side.
Try Foundry Free for 7 DaysNo charge during the trial. Cancel anytime.
Competitor pricing and credit rates verified September 2026 from each provider's published plans. Confirm current rates at each provider's website before purchasing.