Pick Wan if you want cheap, long clips, or a model you can run on your own computer. Wan 2.1 and 2.2 have open weights, so you can download them and run them on your own GPU (graphics card). You can also train a LoRA (a small add-on that learns your lead's face).
The newer versions, 2.5 through 3.0, are API-only: you pay per clip and they run on Alibaba's servers. That split decides how you'll use Wan for a vertical series, so sort it out before you spend anything.
What Wan is good and bad at for micro drama
Use it for:
- Price. Wan 3.0 lists at $0.10 per second at 720p on Alibaba Cloud Model Studio and fal. That's cheap enough to do real retakes.
- Long clips. Wan 3.0 makes clips of 2–30 seconds in one go, with sound. Thirty seconds fits a full beat: the slap, the reaction, the line, the walk-out.
- Several shots in one prompt. Wan 2.6, 2.7 and 3.0 read a timed shot list like "Shot 1 [0–3 s] ...". So one generation can cut from a close-up to a reaction.
- Open weights. Wan 2.1 and 2.2 use the Apache-2.0 license, so you can train a character LoRA on them. A LoRA is one of the most reliable ways to keep a face steady.
Watch out for:
- Dialogue. Alibaba's own prompt guide says lip sync (mouths matching the exact words) "does not work reliably." One creator found Wan 3.0 kept the face better than Seedance 2.5 and cost less, but its dialogue from reference images sounded artificial.
- The open models are small. Wan 2.2 A14B tops out at 720p and makes 81 frames at 16 fps (frames per second) by default, about five seconds. The official README's single-GPU commands for the 14B models need at least 80 GB of VRAM (graphics card memory).
- Text on screen. Contracts, phone screens and DNA results come out unreadable. Add them in the edit.
Versions at a glance
The short version: 2.1 and 2.2 are free to download, and everything from 2.5 up is API-only. In the table, T2V means text-to-video and I2V means image-to-video (you give it a still and it animates it).
| Version | What it's for | Clip length | Weights |
|---|---|---|---|
| Wan 2.1 (1.3B, 14B, FLF2V, VACE) | T2V, I2V, first-last frame, editing | about 5 s | Apache-2.0 |
| Wan 2.2 A14B and TI2V-5B | T2V and I2V; 5B does 720p at 24 fps | 81 frames default | Apache-2.0 |
| Wan 2.2 S2V-14B | Makes a face talk or sing from an image plus audio | extends in chunks | Apache-2.0 |
| Wan 2.2 Animate, Wan-Animate-2 | Moves a still character using a reference performance, or swaps the actor | 2–30 s via API | Apache-2.0 |
| Wan 2.5 preview | First Wan with synced audio | 5 or 10 s | API only |
| Wan 2.6 | Multi-shot, audio, up to 3 reference characters | up to 15 s | API only |
| Wan 2.7 | First-last frame, video continuation, reference, editing | up to 15 s (reference: 10 s) | API only |
| Wan 3.0 and 3.0 Prime | All-in-one: text, first-last frame, up to 10 images, 5 videos and 5 audio refs | 2–30 s | API only |
First-last frame means you give the model the opening and closing image, and it fills in the motion between them. Wan 3.0 opened as a public beta in August 2026, and Prime is the faster, pricier version.
We checked the Wan-AI page on Hugging Face on October 1, 2026. It had repos for 2.1, 2.2, Animate-2 and Dancer, and nothing for 2.5 or later.
Where to run it
- The official app (create.wan.video): The free plan gives you daily check-in credits. App credits don't cover the API.
- The app's paid plans: Pro is $10 a month, or $5 a month billed yearly. It gets you 300 credits, 1080p, 10–30 second videos and no watermark. Premium is $40 a month, or $20 billed yearly, for 1,200 credits.
- Alibaba Cloud Model Studio: This is Alibaba's own API. It's async: you send a job, get a task ID and check back for the video. The Singapore region gives you a free quota for 90 days: 30 seconds for Wan 3.0, 50 seconds for each 2.x model.
- fal and Replicate: They host the same models and bill per second. They're handy if you already use Kling or Seedance there.
- ComfyUI: A free app where you build a video pipeline by wiring boxes (nodes) together. Built-in workflows cover the open models (2.1, 2.2, S2V, Animate, Animate 2). Paid Partner nodes run Wan 2.7 and 3.0 in the same graph.
- Your own GPU: Wan 2.1 1.3B needs 8.19 GB of VRAM. ComfyUI's docs say Wan 2.2 5B "should fit well on 8GB vram" with its built-in offloading. The full 14B models want datacenter cards, or smaller compressed (GGUF) builds made by the community.
For setup help, read our ComfyUI guide and self-hosting your pipeline.
Pricing (as of October 2026)
Most Wan options cost $0.10 a second at 720p, wherever you buy. Prime costs more, and the open 2.2 model on fal costs less.
| Where | Model | 480p | 720p | 1080p |
|---|---|---|---|---|
| Model Studio (Singapore) | wan3.0-video | $0.05/s | $0.10/s | $0.20/s |
| Model Studio (Singapore) | wan3.0-video-prime | $0.068/s | $0.14/s | $0.28/s |
| Model Studio (Singapore) | wan2.7-t2v, wan2.6-t2v | n/a | $0.10/s | $0.15/s |
| fal | Wan 3.0 / 3.0 Prime | $0.05 / $0.068 | $0.10 / $0.14 | $0.20 / $0.28 |
| fal | Wan 2.2 A14B (per video second at 16 fps) | $0.04 | $0.08 | n/a |
| Replicate | wan-2.6-t2v | n/a | $0.10/s | $0.15/s |
| Replicate | wan-2.2-i2v-a14b | $0.40/video | $1.00/video | n/a |
Model Studio shows a "limited-time 30% off" tag on Wan 3.0. Its Global-deployment regions list it lower, at $0.082513/s for 720p. Wan 3.0 also charges for the seconds of any reference video you upload, so check the console before you budget.
Quick math: a 90-second episode at 720p on Wan 3.0 costs $9 if you keep every clip. If you keep one take in three, budget about $27 an episode. Self-hosting drops the per-second fee, but you pay for GPU time and your own hours.
The cost breakdown guide compares Wan with the other four models.
Best practices for 9:16 serialized drama
Lock the cast before episode one
Make a character sheet first: one image showing the face from the front, three-quarter and side. Then animate from that image, not from text alone.
One creator tested Wan 2.6 with the same prompt, with and without his character image. Without the image, he got a different face. With it, the likeness held.
On Wan 3.0 and 2.7, upload the sheet as a reference and call it "Image 1" in the prompt. Images and videos get numbered separately, in upload order.
Paste the exact same character description into every prompt, down to the cufflinks. On the open models, train a LoRA instead.
One tutorial suggests 10 or more clear images, a single trigger word (a made-up word that calls up your character) and about 24 GB of VRAM. On fal, the hosted Wan 2.2 I2V trainer charges $0.005 a step, so 1,000 steps costs $5. Our character consistency guide goes deeper.
Prompt the way Wan reads
Put the camera and light tags first, like this: "Night, practical light, medium close-up, center composition." Alibaba's formula is who's in it, where they are and what moves. Then add light, shot size, lens, angle, tone and style.
For image-to-video, the image already sets the look. So describe only the motion and the camera. Write "fixed camera" when you want the frame to stay still.
Pick shots that read on a phone
Close-ups and medium shots of one person carry vertical drama. Use wide shots only to show a location, and keep them short. Over-the-shoulder shots work for fights if you say who's in the foreground.
Insert shots of props (a ring box, a cracked phone, a wolf-claw scratch) are cheap. They also let you cut around weak takes.
Handle dialogue in layers
On 2.5 and later, write one line per clip. Tie it to an action and a voice label: Elena slams the folder down. [Elena, shaky but defiant]: "I quit."
Alibaba's guide recommends a unique label for each character. It also suggests linking words like "Immediately," so two speakers don't blur together.
Need the exact words? Record or generate the line first, then drive the clip with it. Wan 2.7 image-to-video accepts driving audio, and the open S2V model animates a face from an image plus audio.
Expect to fix some mouths afterward. See lip sync and dubbing.
Stitch clips on purpose
Use the last frame of one clip as the first frame of the next. Wan 2.7 and 3.0 support this with first-last frame, and 2.7 can also continue a video. For one continuous take on 2.7, write "Generate single shot."
Mixing works too. One creator cut a 15-second multi-shot clip together with a single-shot one and kept only the strong parts.
Pace for the scroll
Start every clip mid-action, because the first three seconds of your episode are the hook. In a multi-shot prompt, give each shot 2–5 seconds. End on a reaction face before the cliffhanger.
Common pitfalls
- Prompt rewriting is on by default. It helps short prompts but can change your blocking (who stands where). Turning it off saves 20–60 seconds and may cost some quality, according to fal, so test both.
- Wrong aspect ratio. On fal, Wan 2.6 and 2.7 default to 16:9, and 3.0 defaults to "adaptive." Set 9:16 every time (720×1280 or 1080×1920).
- Scene changes inside one shot. Alibaba's guide warns that cuts happen between clips, not within them. Use timestamps or separate clips.
- Real names. Prompts that name real people get rejected or drift (the face slowly changes). Describe the person instead.
- 16 fps on the open models. Slow camera moves can look choppy. Interpolate them (add in-between frames), since fal's Wan 2.2 endpoints offer FILM or RIFE for this.
- Look-alike extras. In one creator's Wan 2.6 tests, a group of runners came out looking like brothers. Give background characters different hair, props or clothes.
Prompts to start with
Start with our Wan prompt library: 40 original prompts with settings for each version. Or jump to a genre: CEO romance, revenge, werewolf, historical, thriller, comedy, fantasy or family.
For node graphs, check the open-source workflow projects and model tooling. Still deciding? Read choosing a video model.
Sources
- github.com/Wan-Video/Wan2.1
- github.com/Wan-Video/Wan2.2
- github.com/Wan-Video/Wan-Animate-2
- huggingface.co/Wan-AI
- www.alibabacloud.com/help/en/model-studio/video-generation
- www.alibabacloud.com/help/en/model-studio/model-pricing
- www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference
- www.alibabacloud.com/help/en/model-studio/text-to-video-prompt
- modelstudio.console.alibabacloud.com/model-releases/wan3.0-video
- wan.video/pricing
- fal.ai/models/alibaba/wan-3.0/reference-to-video
- fal.ai/models/alibaba/wan-3.0-prime/text-to-video
- fal.ai/models/fal-ai/wan/v2.7/text-to-video
- fal.ai/models/wan/v2.6/text-to-video
- fal.ai/models/fal-ai/wan/v2.2-a14b/text-to-video
- fal.ai/models/fal-ai/wan/v2.2-14b/speech-to-video
- fal.ai/models/fal-ai/wan-22-trainer/i2v-a14b
- replicate.com/wan-video/wan-2.6-t2v
- replicate.com/wan-video/wan-2.2-i2v-a14b
- docs.comfy.org/tutorials/video/wan/wan2_2
- docs.comfy.org/built-in-nodes/Wan3ReferenceToVideoApi
- ourcodeworld.com/articles/read/4473/wan-3-0-hands-on-with-alibaba-s-ai-video-model-generating-30-second-clips
- www.youtube.com/watch?v=HjFX9c-wRJc
- www.youtube.com/watch?v=1SLe6kKR41g
- www.youtube.com/watch?v=HlXmji1O_bI
- www.youtube.com/watch?v=DcmTAaKgGXM
- www.youtube.com/watch?v=vvHkFbjJf8A