Skip to content

Choosing a Video Model for AI Micro Drama: Seedance, Kling, Veo, Hailuo or Wan

Pick Seedance, Kling, Veo, Hailuo or Wan for your vertical drama: a side-by-side table of clip length, audio, references and price, plus a quick test.

Updated Oct 1, 2026

Pick the model that fixes the shot you're stuck on, not the "best" model overall. A face-to-face fight with three lines of dialogue needs different things than a wide shot of a mansion. A team making 100 episodes also cares about the per-second price more than a hobbyist making one pilot.

Below, we compare the five model families we cover, as of October 1, 2026. Then there's a quick test to run before you commit a whole season to one of them.

The short version

  • Dialogue-heavy scenes and long takes: Seedance 2.5. Up to 30 seconds per clip, with lip-synced speech and lots of reference images.
  • Emotional close-ups on a budget: Kling 3.0. It makes sound, can cut several shots from one prompt, and the official API starts at $0.112 a second.
  • Polished English dialogue, in Google's tools: Veo 3.1 (and Google's newer Gemini Omni Flash). Short clips and strong audio, easy to reach through Flow.
  • Lowest price per second with sound: MiniMax H3 from the Hailuo line, at $0.06 a second at 768p on fal.
  • Self-hosting, or long takes on the cheap: Wan. The 2.x models are open weights under Apache-2.0, and Wan 3.0 runs up to 30 seconds through the API.

Two terms you'll see a lot: lip sync means the mouth movements match the words. Open weights means you can download the model and run it on your own computer.

Side by side

Seedance 2.5Kling 3.0Veo 3.1MiniMax H3 (Hailuo)Wan 3.0
VendorByteDanceKuaishouGoogleMiniMaxAlibaba
Clip length per generation4–30s3–15s4, 6 or 8s; extend +7s per step4–15s2–30s
9:16 outputYesYesYesYesYes
Native audio and lip syncYesYesYes, always on in the Gemini APIYes, stereoYes
Dialogue languages (vendor-listed)11, including English, Spanish, PortugueseChinese, English, Japanese, Korean, SpanishEnglish fully supported; others not evaluated11, including English, French, GermanCheck the official docs
Character referencesUp to 30 images, 10 videos, 10 audioElements: frontal image plus up to 3 references, optional voiceUp to 3 reference images (8s clips only)Up to 9 images, 3 videos, 3 audioUp to 10 images, 5 videos, 5 audio
Multi-shot in one promptYes (Shot labels or timestamps)Yes, up to 6 shots in 15sYes, timestamp promptingYesYes
Open weightsNoNoNoYes, community license with territory limitsNo (Wan 2.1 and 2.2 are, Apache-2.0)
Lowest listed price, 720p with audioabout $0.23/s (BytePlus, Replicate)$0.112/s (3.0 Turbo, official API)$0.40/s (Fast $0.10, Lite $0.05)$0.06/s at 768p (fal)$0.10/s
Our full guideSeedanceKlingVeoHailuoWan

The takeaway: all five make vertical video with sound, so the real differences are clip length, reference limits and price.

A few things the table can't show:

  • Kling 4.0 is coming. Kuaishou's release page says it "will officially launch in October," with 3–30 second clips and up to 10 reference images. A Flash version is in early access for a small group, but Kling's prepaid API packages don't cover 4.0 yet.
  • Google now points developers to Gemini Omni Flash as its default video model. It keeps Veo 3.1 for extending scenes, last-frame control and older pipelines. Omni Flash costs about $0.10 a second at 720p in the Gemini API.
  • MiniMax H3's open-weight license doesn't cover the EU, the UK, South Korea or the USA. If you're there, you need a formal license from MiniMax to run it yourself. Using it through the API is a separate matter.
  • Sora is gone. OpenAI's help center says the Sora app shut down on April 26, 2026, and the API on September 24, 2026. Older tutorials that recommend it are out of date.

Pick by your problem

"My dialogue scenes look fake"

Try Seedance 2.5, Kling 3.0 and Veo 3.1 first. All three make speech together with the picture.

Creators testing these models report swapped or garbled lines when two people talk in one clip. So give each clip one speaker, and cut between them as shot and reverse shot.

Happy with a silent model's look? You can add voices afterwards with a lip-sync pass. The lip sync and dubbing guide compares both routes.

"My lead looks different every episode"

Here, reference support matters more than raw image quality. A reference is an image you give the model so it copies that face.

Seedance 2.5, MiniMax H3 and Wan 3.0 take the most reference files, and Kling's Elements are built for exactly this. Veo caps you at three reference images, which is fine for two characters but tight for a big cast.

Whatever you pick, build character sheets before episode 1. The character consistency guide walks you through it.

"I can't afford 100 episodes"

Start with the budget models: MiniMax H3, Veo 3.1 Lite and Seedance 2.0 Mini. Kling 3.0 Turbo sits in the middle, with audio included.

In the table above, the lowest listed price ($0.06 a second) and the highest ($0.40) are nearly seven times apart. The cost breakdown turns these prices into per-minute and per-season numbers.

"I want my scene in one long take"

Use Seedance 2.5 or Wan 3.0. Both go up to 30 seconds in a single clip.

Veo can go longer by chaining 7-second extensions. But Google's docs limit extensions to 720p, and creators report faces changing over long chains.

"I need camera control"

The Hailuo line gives you the most direct camera commands. On 2.3 you write bracket commands like Push in, Truck left and Tracking shot. On H3 you describe the motion, its size and its speed.

Kling and Veo also follow plain camera language well. See vertical framing and shot language for which moves work in 9:16.

"I don't want per-second fees at all"

Run Wan 2.2 yourself. The models are Apache-2.0 on GitHub and Hugging Face, and ComfyUI supports them out of the box. ComfyUI is a free app for building AI video workflows.

You swap fees for GPU time and setup work.

The 14B models need serious hardware. The README's single-GPU commands call for at least 80 GB of VRAM (your graphics card's memory). The 5B model fits normal gaming cards if you let it offload to system memory.

Start with ComfyUI for micro drama and the self-hosting guide.

"I want to use a real actor's face"

Check each platform's rules before you build around it. BytePlus doesn't accept reference images or videos of real human faces for Seedance, except through its verified-asset routes. Veo's image-based modes only allow adults.

Get written consent from anyone whose face or voice you use. Then read the copyright basics guide.

Run this test before you commit

Demo reels won't show you how a model handles your story. Spend an hour and a few dollars on this test instead:

  1. Pick three shots from your episode 1: one dialogue close-up, one two-person medium shot and one wide establishing shot. Each of our script templates has a shot list you can borrow.
  2. Use the same character sheet and shot descriptions on every model. Our prompt library has the same scenes written for each model: Seedance, Kling, Veo, Hailuo, Wan.
  3. Make each shot four times at 9:16 and your target resolution.
  4. Count keepers, not the best take. How many of the four would you actually use? Check the face, lip sync, hands, unwanted subtitles and whether the camera did what you asked.
  5. Do the math. Your keep rate tells you how many takes you really need per shot. That number decides your budget.

You'll probably end up with two models. One handles faces and dialogue, and a cheaper one handles wide shots and inserts.

That's normal. Just keep each character's close-ups on the same model so their face doesn't change mid-season, and match the colors in the edit.

Related

Sources

  1. seed.bytedance.com/en/seedance2_5
  2. docs.byteplus.com/en/docs/ModelArk/1099320
  3. docs.byteplus.com/en/docs/ModelArk/2607688
  4. docs.byteplus.com/en/docs/ModelArk/2608626
  5. kling.ai/quickstart/klingai-video-3-model-user-guide
  6. kling.ai/dev/pricing
  7. kling.ai/dev/model-release/kling-4
  8. ai.google.dev/gemini-api/docs/veo
  9. ai.google.dev/gemini-api/docs/video
  10. ai.google.dev/gemini-api/docs/pricing
  11. platform.minimax.io/docs/guides/video-generation
  12. platform.minimax.io/docs/guides/pricing-paygo
  13. huggingface.co/MiniMaxAI/MiniMax-H3
  14. huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
  15. fal.ai/models/minimax/h3/text-to-video
  16. www.alibabacloud.com/help/en/model-studio/video-generation
  17. www.alibabacloud.com/help/en/model-studio/model-pricing
  18. github.com/Wan-Video/Wan2.2
  19. help.openai.com/en/articles/20001152-what-to-know-about-the-sora-discontinuation