Skip to content

Wan for AI Micro Drama: Versions, Pricing and 9:16 Prompting

Wan for AI micro drama: free-to-download 2.2 vs API-only 3.0, real per-second prices, 9:16 settings, and how to keep faces the same across episodes.

Alibaba (Tongyi Lab) · updated Oct 1, 2026

Pick Wan if you want cheap, long clips, or a model you can run on your own computer. Wan 2.1 and 2.2 have open weights, so you can download them and run them on your own GPU (graphics card). You can also train a LoRA (a small add-on that learns your lead's face).

The newer versions, 2.5 through 3.0, are API-only: you pay per clip and they run on Alibaba's servers. That split decides how you'll use Wan for a vertical series, so sort it out before you spend anything.

What Wan is good and bad at for micro drama

Use it for:

  • Price. Wan 3.0 lists at $0.10 per second at 720p on Alibaba Cloud Model Studio and fal. That's cheap enough to do real retakes.
  • Long clips. Wan 3.0 makes clips of 2–30 seconds in one go, with sound. Thirty seconds fits a full beat: the slap, the reaction, the line, the walk-out.
  • Several shots in one prompt. Wan 2.6, 2.7 and 3.0 read a timed shot list like "Shot 1 [0–3 s] ...". So one generation can cut from a close-up to a reaction.
  • Open weights. Wan 2.1 and 2.2 use the Apache-2.0 license, so you can train a character LoRA on them. A LoRA is one of the most reliable ways to keep a face steady.

Watch out for:

  • Dialogue. Alibaba's own prompt guide says lip sync (mouths matching the exact words) "does not work reliably." One creator found Wan 3.0 kept the face better than Seedance 2.5 and cost less, but its dialogue from reference images sounded artificial.
  • The open models are small. Wan 2.2 A14B tops out at 720p and makes 81 frames at 16 fps (frames per second) by default, about five seconds. The official README's single-GPU commands for the 14B models need at least 80 GB of VRAM (graphics card memory).
  • Text on screen. Contracts, phone screens and DNA results come out unreadable. Add them in the edit.

Versions at a glance

The short version: 2.1 and 2.2 are free to download, and everything from 2.5 up is API-only. In the table, T2V means text-to-video and I2V means image-to-video (you give it a still and it animates it).

VersionWhat it's forClip lengthWeights
Wan 2.1 (1.3B, 14B, FLF2V, VACE)T2V, I2V, first-last frame, editingabout 5 sApache-2.0
Wan 2.2 A14B and TI2V-5BT2V and I2V; 5B does 720p at 24 fps81 frames defaultApache-2.0
Wan 2.2 S2V-14BMakes a face talk or sing from an image plus audioextends in chunksApache-2.0
Wan 2.2 Animate, Wan-Animate-2Moves a still character using a reference performance, or swaps the actor2–30 s via APIApache-2.0
Wan 2.5 previewFirst Wan with synced audio5 or 10 sAPI only
Wan 2.6Multi-shot, audio, up to 3 reference charactersup to 15 sAPI only
Wan 2.7First-last frame, video continuation, reference, editingup to 15 s (reference: 10 s)API only
Wan 3.0 and 3.0 PrimeAll-in-one: text, first-last frame, up to 10 images, 5 videos and 5 audio refs2–30 sAPI only

First-last frame means you give the model the opening and closing image, and it fills in the motion between them. Wan 3.0 opened as a public beta in August 2026, and Prime is the faster, pricier version.

We checked the Wan-AI page on Hugging Face on October 1, 2026. It had repos for 2.1, 2.2, Animate-2 and Dancer, and nothing for 2.5 or later.

Where to run it

  • The official app (create.wan.video): The free plan gives you daily check-in credits. App credits don't cover the API.
  • The app's paid plans: Pro is $10 a month, or $5 a month billed yearly. It gets you 300 credits, 1080p, 10–30 second videos and no watermark. Premium is $40 a month, or $20 billed yearly, for 1,200 credits.
  • Alibaba Cloud Model Studio: This is Alibaba's own API. It's async: you send a job, get a task ID and check back for the video. The Singapore region gives you a free quota for 90 days: 30 seconds for Wan 3.0, 50 seconds for each 2.x model.
  • fal and Replicate: They host the same models and bill per second. They're handy if you already use Kling or Seedance there.
  • ComfyUI: A free app where you build a video pipeline by wiring boxes (nodes) together. Built-in workflows cover the open models (2.1, 2.2, S2V, Animate, Animate 2). Paid Partner nodes run Wan 2.7 and 3.0 in the same graph.
  • Your own GPU: Wan 2.1 1.3B needs 8.19 GB of VRAM. ComfyUI's docs say Wan 2.2 5B "should fit well on 8GB vram" with its built-in offloading. The full 14B models want datacenter cards, or smaller compressed (GGUF) builds made by the community.

For setup help, read our ComfyUI guide and self-hosting your pipeline.

Pricing (as of October 2026)

Most Wan options cost $0.10 a second at 720p, wherever you buy. Prime costs more, and the open 2.2 model on fal costs less.

WhereModel480p720p1080p
Model Studio (Singapore)wan3.0-video$0.05/s$0.10/s$0.20/s
Model Studio (Singapore)wan3.0-video-prime$0.068/s$0.14/s$0.28/s
Model Studio (Singapore)wan2.7-t2v, wan2.6-t2vn/a$0.10/s$0.15/s
falWan 3.0 / 3.0 Prime$0.05 / $0.068$0.10 / $0.14$0.20 / $0.28
falWan 2.2 A14B (per video second at 16 fps)$0.04$0.08n/a
Replicatewan-2.6-t2vn/a$0.10/s$0.15/s
Replicatewan-2.2-i2v-a14b$0.40/video$1.00/videon/a

Model Studio shows a "limited-time 30% off" tag on Wan 3.0. Its Global-deployment regions list it lower, at $0.082513/s for 720p. Wan 3.0 also charges for the seconds of any reference video you upload, so check the console before you budget.

Quick math: a 90-second episode at 720p on Wan 3.0 costs $9 if you keep every clip. If you keep one take in three, budget about $27 an episode. Self-hosting drops the per-second fee, but you pay for GPU time and your own hours.

The cost breakdown guide compares Wan with the other four models.

Best practices for 9:16 serialized drama

Lock the cast before episode one

Make a character sheet first: one image showing the face from the front, three-quarter and side. Then animate from that image, not from text alone.

One creator tested Wan 2.6 with the same prompt, with and without his character image. Without the image, he got a different face. With it, the likeness held.

On Wan 3.0 and 2.7, upload the sheet as a reference and call it "Image 1" in the prompt. Images and videos get numbered separately, in upload order.

Paste the exact same character description into every prompt, down to the cufflinks. On the open models, train a LoRA instead.

One tutorial suggests 10 or more clear images, a single trigger word (a made-up word that calls up your character) and about 24 GB of VRAM. On fal, the hosted Wan 2.2 I2V trainer charges $0.005 a step, so 1,000 steps costs $5. Our character consistency guide goes deeper.

Prompt the way Wan reads

Put the camera and light tags first, like this: "Night, practical light, medium close-up, center composition." Alibaba's formula is who's in it, where they are and what moves. Then add light, shot size, lens, angle, tone and style.

For image-to-video, the image already sets the look. So describe only the motion and the camera. Write "fixed camera" when you want the frame to stay still.

Pick shots that read on a phone

Close-ups and medium shots of one person carry vertical drama. Use wide shots only to show a location, and keep them short. Over-the-shoulder shots work for fights if you say who's in the foreground.

Insert shots of props (a ring box, a cracked phone, a wolf-claw scratch) are cheap. They also let you cut around weak takes.

Handle dialogue in layers

On 2.5 and later, write one line per clip. Tie it to an action and a voice label: Elena slams the folder down. [Elena, shaky but defiant]: "I quit."

Alibaba's guide recommends a unique label for each character. It also suggests linking words like "Immediately," so two speakers don't blur together.

Need the exact words? Record or generate the line first, then drive the clip with it. Wan 2.7 image-to-video accepts driving audio, and the open S2V model animates a face from an image plus audio.

Expect to fix some mouths afterward. See lip sync and dubbing.

Stitch clips on purpose

Use the last frame of one clip as the first frame of the next. Wan 2.7 and 3.0 support this with first-last frame, and 2.7 can also continue a video. For one continuous take on 2.7, write "Generate single shot."

Mixing works too. One creator cut a 15-second multi-shot clip together with a single-shot one and kept only the strong parts.

Pace for the scroll

Start every clip mid-action, because the first three seconds of your episode are the hook. In a multi-shot prompt, give each shot 2–5 seconds. End on a reaction face before the cliffhanger.

Common pitfalls

  • Prompt rewriting is on by default. It helps short prompts but can change your blocking (who stands where). Turning it off saves 20–60 seconds and may cost some quality, according to fal, so test both.
  • Wrong aspect ratio. On fal, Wan 2.6 and 2.7 default to 16:9, and 3.0 defaults to "adaptive." Set 9:16 every time (720×1280 or 1080×1920).
  • Scene changes inside one shot. Alibaba's guide warns that cuts happen between clips, not within them. Use timestamps or separate clips.
  • Real names. Prompts that name real people get rejected or drift (the face slowly changes). Describe the person instead.
  • 16 fps on the open models. Slow camera moves can look choppy. Interpolate them (add in-between frames), since fal's Wan 2.2 endpoints offer FILM or RIFE for this.
  • Look-alike extras. In one creator's Wan 2.6 tests, a group of runners came out looking like brothers. Give background characters different hair, props or clothes.

Prompts to start with

Start with our Wan prompt library: 40 original prompts with settings for each version. Or jump to a genre: CEO romance, revenge, werewolf, historical, thriller, comedy, fantasy or family.

For node graphs, check the open-source workflow projects and model tooling. Still deciding? Read choosing a video model.

Sources

  1. github.com/Wan-Video/Wan2.1
  2. github.com/Wan-Video/Wan2.2
  3. github.com/Wan-Video/Wan-Animate-2
  4. huggingface.co/Wan-AI
  5. www.alibabacloud.com/help/en/model-studio/video-generation
  6. www.alibabacloud.com/help/en/model-studio/model-pricing
  7. www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference
  8. www.alibabacloud.com/help/en/model-studio/text-to-video-prompt
  9. modelstudio.console.alibabacloud.com/model-releases/wan3.0-video
  10. wan.video/pricing
  11. fal.ai/models/alibaba/wan-3.0/reference-to-video
  12. fal.ai/models/alibaba/wan-3.0-prime/text-to-video
  13. fal.ai/models/fal-ai/wan/v2.7/text-to-video
  14. fal.ai/models/wan/v2.6/text-to-video
  15. fal.ai/models/fal-ai/wan/v2.2-a14b/text-to-video
  16. fal.ai/models/fal-ai/wan/v2.2-14b/speech-to-video
  17. fal.ai/models/fal-ai/wan-22-trainer/i2v-a14b
  18. replicate.com/wan-video/wan-2.6-t2v
  19. replicate.com/wan-video/wan-2.2-i2v-a14b
  20. docs.comfy.org/tutorials/video/wan/wan2_2
  21. docs.comfy.org/built-in-nodes/Wan3ReferenceToVideoApi
  22. ourcodeworld.com/articles/read/4473/wan-3-0-hands-on-with-alibaba-s-ai-video-model-generating-30-second-clips
  23. www.youtube.com/watch?v=HjFX9c-wRJc
  24. www.youtube.com/watch?v=1SLe6kKR41g
  25. www.youtube.com/watch?v=HlXmji1O_bI
  26. www.youtube.com/watch?v=DcmTAaKgGXM
  27. www.youtube.com/watch?v=vvHkFbjJf8A

Open-source projects that support Wan

Writes a narrated short video from a topic: script, voiceover, footage, subtitles

Video framework
  • MiniMax H3
  • Seedance
  • Wan
  • +7
128k GitHub starstodayMIT

calesthio

OpenMontage

Video studio that your AI coding assistant runs

Video framework
  • Google Veo
  • Kling
  • Seedance 2.0
  • +7
62k GitHub stars25 days agoAGPL-3.0

Type one topic and get a narrated short video: script, images, voice, music

Video framework
  • Wan 2.1
  • DashScope Wan
  • HappyHorse
  • +7
29k GitHub stars3 months agoApache-2.0

Wan-Video

Wan2.2

Open video models for text, image, speech and character animation

Model tooling
  • Wan 2.2-T2V-A14B
  • Wan 2.2-I2V-A14B
  • Wan 2.2-TI2V-5B
  • +5
18k GitHub stars10 days agoApache-2.0

chatfire-AI

huobao-drama

Novel in, finished short-drama episode out, with script and storyboard

Platform
  • Seedance 2.0
  • MiniMax H3
  • Wan 3.0
  • +2
16k GitHub stars3 days agoCC-BY-NC-SA-4.0

Trains and runs image, video and audio diffusion models on your own GPU

Model tooling
  • MiniMax H3
  • LTX-2.5
  • Wan-Animate-2
  • +7
13k GitHub starsyesterdayApache-2.0