Pick one of two ways to make your characters talk. You can let the video model make the voice and the picture together. Or you can make the voice first and lip-sync it onto the picture later (lip sync means mouth movements that match the words).
Most serious productions end up using both. Dialogue is where AI video breaks fastest: mouths chewing the air, voices that change between shots, two characters talking with one face.
We checked the prices below on October 1, 2026. They change often, so confirm them on the linked pages.
Route 1: models that make speech with the picture
Several current models output dialogue, sound effects and background sound in the same pass as the video. You write the line in the prompt, and the mouth moves to it. Use this table to see what each one supports.
| Model | Native speech | Language notes from the docs | Clip length per generation |
|---|---|---|---|
| Veo 3.1 | Yes, audio is always on in the Gemini API | English fully supported, other languages not evaluated | 4, 6 or 8 seconds, extendable |
| Seedance 2.5 | Yes, toggle on fal (same price either way) | Chinese, English, Spanish, Japanese, Korean, Portuguese, Arabic and several Southeast Asian languages | 4 to 30 seconds |
| Kling 3.0 | Yes, plus optional voice control | Chinese, English, Japanese, Korean, Spanish; other languages get translated to English | 3 to 15 seconds |
| MiniMax H3 | Yes, stereo | Dialogue described as stable in 11 languages | 4 to 15 seconds |
| Wan 2.6 / 2.7 / 3.0 | Yes, audio sync (API only) | Check the model page per version | Up to 15 seconds on 2.6/2.7, up to 30 on 3.0 |
| LTX-2.x (open weights) | Yes, audio and video in one model | See the repo | See the repo |
The takeaway: all the big models can talk now, but language support and clip length differ a lot.
Our model guides go deeper on each one: Veo, Seedance, Kling, Hailuo and MiniMax H3 and Wan.
What creators keep learning the hard way
Use one speaker per clip. When two characters share a talky clip, models swap lines, repeat them, or move both mouths. One Veo creator found a dull but reliable fix: give each character's line its own clip, then cut them together.
A creator using MiniMax H3 said the same thing another way. Let one character speak while the other keeps their mouth closed. And test a 6-second version before you pay for a long clip.
Keep lines short. A 5 to 8 second clip holds one line easily. If the line doesn't fit, it's probably two shots anyway.
Write dialogue the way each model reads it. Veo's docs put dialogue in quotation marks. fal's Kling 3.0 page says to write English speech in lowercase, with capitals only for acronyms and names. Seedance's prompt guide asks you to name the language before any dialogue that isn't Chinese.
Our prompt library shows each format in real prompts.
Watch for voice drift. The same character can sound like three different people across one episode. Use the model's voice reference features where they exist, like Kling's voice IDs or audio references on Seedance and H3.
If your model doesn't have one, run every line for a character through one voice changer at the end so they match.
Watch for burned-in subtitles. Models sometimes add captions you didn't ask for. Seedance's guide says it happens more in vertical video.
Add "no subtitles" to the prompt. You'll add your own captions later anyway. See editing and captions.
Route 2: voice first, then lip-sync
This way, you lock the performance in audio first, then make the picture match. It's slower. But you control every word, pause and retake, and you can reuse the same cloned voice for 80 episodes.
The steps:
- Write the line and make it with TTS (text-to-speech, a tool that reads text aloud in a chosen voice). Or record a human actor.
- Make the shot silent or with background sound only. Or start from a still frame.
- Run a lip-sync model to move the mouth to your audio.
There are two kinds of lip-sync tools. Video-to-video tools re-animate the mouth on a clip you already have. Audio-driven image-to-video tools take a still portrait plus an audio file and make a talking shot, often with head and body movement.
This table compares the prices.
| Tool (via fal unless noted) | Type | Listed price |
|---|---|---|
| Kling LipSync | Video-to-video | $0.014 per input video second, rounded up to 5-second increments |
| sync. Lipsync 2.0 | Video-to-video | $3 per minute of video |
| LatentSync | Video-to-video | $0.20 for videos up to 40 seconds, then $0.005 per second |
| VEED Lipsync v2 | Video-to-video | $0.07 per output second |
| MiniMax H3 Max Lip Sync | Image plus audio to video | $0.05/s at 480p, $0.08/s at 768p, $0.16/s at 1080p |
| Wan 2.2 Speech-to-Video | Image plus audio to video | $0.10, $0.15 or $0.20 per second at 480p, 580p or 720p |
The takeaway: for a short dialogue clip of 5 to 10 seconds, Kling LipSync costs the least of the video-to-video tools here.
An ElevenLabs walkthrough compared audio-driven models. Wan 2.6 gave the sharpest picture but moved the camera and body a lot, so the presenter added "still continuous shot" to the prompt.
In the same video, OmniHuman's clip length followed the uploaded audio, up to 30 seconds. Creatify Aurora allowed longer clips but topped out at 720p, so it needed an upscale (enlarging the video with an AI tool). Expect every tool to have one strength and one annoying habit.
Voices: TTS and cloning costs
Voice cloning means making a copy of a real voice so the tool can speak new lines in it. Here's what voices cost.
| Service | What you pay | Notes |
|---|---|---|
| ElevenLabs Free | $0, 10k credits a month | No commercial license on the free plan |
| ElevenLabs Starter | $6/month, 30k credits | Adds commercial license and Instant Voice Cloning |
| ElevenLabs Creator | $22/month ($11 first month), 121k credits | Adds Professional Voice Cloning |
| MiniMax Speech 2.8 | $100 per million characters (HD), $60 (Turbo) | Rapid voice cloning $1.5 per voice, voice design $3 per voice |
| Chatterbox on fal | $0.025 per 1,000 characters | Open model, hosted |
The takeaway: voices are cheap, and you need a paid plan before you can use them commercially on ElevenLabs.
ElevenLabs says text-to-speech costs about 1 credit per character on its v2 Multilingual models. Its dubbing products cost 2,000 to 10,000 credits per minute, depending on mode and watermark.
Here's the rough math. An episode with around 1,000 characters of dialogue uses about 1,000 credits, so Starter covers a few dozen episodes before retakes. Retakes are where the credits really go.
Open-source options
These free projects cover lip sync and voices. Check the license column before you use any of them in a paid project.
| Project | License | What it's for |
|---|---|---|
| LatentSync | Apache-2.0 | Video-to-video lip-sync. README lists 8 GB VRAM minimum for v1.5, 18 GB for v1.6 |
| MuseTalk | MIT (code) | Real-time lip-sync by latent inpainting; check each bundled model's license |
| InfiniteTalk | Apache-2.0 | Audio-driven dubbing with head and body motion, long clips |
| Wan2.2-S2V | Apache-2.0 | Speech-to-video from a still and an audio track |
| Wav2Lip | Research only | The README says commercial use of the open model is prohibited |
| Fun-CosyVoice 3 | Apache-2.0 | Zero-shot multilingual TTS, 9 languages plus Chinese dialects |
| Chatterbox | MIT | Multilingual TTS with voice cloning |
| F5-TTS | Code MIT, weights CC-BY-NC-4.0 | The published weights are non-commercial |
(VRAM is your graphics card's memory.)
Read the license on the weights, not only the code. The weights are the model file you download. A repo can be MIT while the model file is non-commercial, and the model file is what ends up in your episode.
Building a full setup on your own computer? Start with self-hosting an open-source pipeline. You can browse lip-sync and dubbing repos under editing projects.
Dubbing into other languages
Vertical drama travels well. The same 80 episodes can run in English, Spanish, Portuguese and more. Here's how to dub them:
- Transcribe your final mix (Whisper works fine), and export the dialogue track on its own.
- Translate it with an LLM (a chatbot like ChatGPT or Claude). Then have a fluent speaker fix the slang and the emotional beats. A cliffhanger line that falls flat in translation kills the episode.
- Make the new audio with the same cloned voice for each character, with that actor's written consent.
- Re-run lip sync on dialogue shots. Wide shots and over-the-shoulder shots often don't need it.
- Re-time the captions. Translated lines usually run longer, so check reading speed on a phone.
One creator showed a free route. He used a "JustDubIt" LoRA on LTX-2, run through Wan2GP, that takes an English clip and outputs the same clip speaking Japanese, with matching mouth shapes. (A LoRA is a small add-on trained for one job.)
Under the LTX-2.x Community License, organizations with $10M or more in yearly revenue need a paid license.
Consent isn't optional
Get a signed voice release from every actor whose voice you clone. ElevenLabs' Prohibited Use Policy bans copying someone's voice without consent or legal right, and other providers have similar rules.
Keep the original recordings. Never clone a celebrity, or any real person who hasn't agreed. Our copyright basics guide covers likeness and voice rights in more detail.
Lip-sync QA checklist
- Watch every dialogue shot full size on a phone. Problems that hide on a monitor show up on a 6-inch vertical screen.
- Watch the mouth close on "p", "b" and "m". Those sounds are the first to look wrong.
- Check teeth and tongue for smearing in close-ups.
- Make sure characters who are listening, or off screen, keep their mouths still.
- Play the episode with the sound off, then on. The mouth should still look like speech with no audio.
- Match loudness from shot to shot before you add music. See music and sound.
Want the total cost once dialogue is in? The cost breakdown does the per-minute math. /tools lists the paid voice and lip-sync apps we track.
Sources
- ai.google.dev/gemini-api/docs/veo
- fal.ai/models/bytedance/seedance-2.5/text-to-video
- docs.byteplus.com/en/docs/ModelArk/2607689
- kling.ai/quickstart/klingai-video-3-model-user-guide
- fal.ai/models/fal-ai/kling-video/v3/pro/text-to-video
- huggingface.co/MiniMaxAI/MiniMax-H3
- fal.ai/models/wan/v2.6/text-to-video
- github.com/Wan-Video/Wan2.2
- github.com/Lightricks/LTX-2
- fal.ai/models/fal-ai/kling-video/lipsync/audio-to-video
- fal.ai/models/fal-ai/sync-lipsync/v2
- fal.ai/models/fal-ai/latentsync
- fal.ai/models/veed/lipsync/v2
- fal.ai/models/minimax/h3-max/lip-sync/image-to-video
- fal.ai/models/fal-ai/wan/v2.2-14b/speech-to-video
- elevenlabs.io/pricing
- elevenlabs.io/use-policy
- platform.minimax.io/docs/guides/pricing-paygo
- fal.ai/models/fal-ai/chatterbox/text-to-speech
- github.com/bytedance/LatentSync
- github.com/TMElyralab/MuseTalk
- github.com/MeiGen-AI/InfiniteTalk
- github.com/Rudrabha/Wav2Lip
- github.com/QwenAudio/CosyVoice
- github.com/resemble-ai/chatterbox
- huggingface.co/SWivid/F5-TTS
- www.youtube.com/watch?v=FI2s06CTvHg
- www.youtube.com/watch?v=yDUqZzAQPzI
- www.youtube.com/watch?v=VYH-zwR8iQU
- www.youtube.com/watch?v=Z8giDB2wzFc
- www.youtube.com/watch?v=PeyVgAQ_oxM