Learn five ComfyUI workflows and run them hundreds of times. You don't need one magic setup. A keyframe becomes a clip, two frames become a move, and a still plus a voice line becomes a talking shot.
ComfyUI is a free app for running AI image and video models on your own computer. You build each workflow by linking boxes called nodes, like a flowchart. We use the official templates wherever we can.
Haven't picked your tools yet? Read self-hosting an open-source pipeline first. For prompting tips, see our Wan guide and the Wan prompt library.
Setup
ComfyUI runs on Windows, Linux and macOS (Apple Silicon). The docs offer Comfy Desktop, a Windows portable build and a manual install. They currently recommend Python 3.13.
Video models are much heavier than image models. As one tutorial put it, a 5-second clip at 30 fps is basically 150 images. So you'll need plenty of VRAM (your graphics card's memory) and some patience.
The official templates live in the Template Library inside the app. If a template below is missing, update ComfyUI.
The templates worth knowing
These are the official templates that matter for drama. I2V means image-to-video, T2V means text-to-video, and FLF2V means first-last frame to video.
| Template (from ComfyUI docs) | What it does for a drama | Notes from the docs |
|---|---|---|
| Wan 2.2 5B TI2V | Fast text- or image-to-video drafts at 720p | "Should fit well on 8GB vram" with native offloading |
| Wan 2.2 14B T2V / I2V | Final-quality shots | High-noise and low-noise expert models load as a pair |
| Wan 2.2 14B FLF2V | First-last frame: you set start and end, it fills the move | Uses the same model files as I2V |
| Wan 2.2 S2V | A still plus an audio line becomes a talking shot | 16 fps model; each extend chunk adds 77 frames, about 4.8 seconds |
| Wan 2.2 Animate | Mix mode swaps a character into footage; Move mode drives a character with a performer's motion | Start small in size the first time to avoid running out of VRAM |
| LTX-2 T2V / I2V | Video and audio in one pass, including dialogue and SFX | 19B model; distilled version runs in 8 steps |
| HunyuanVideo 1.5 T2V / I2V | 5 to 10 second clips on consumer GPUs | Docs cite 24 GB VRAM |
The takeaway: draft with Wan 2.2 5B, finish with Wan 2.2 14B, and use S2V or LTX-2 when someone talks.
Kijai's WanVideoWrapper is another way to run Wan. It often gets new research features first. Start with the native templates, and switch only if you need something just the wrapper has.
Recipe 1: keyframe to clip (your workhorse)
Start most shots from a still you've approved, not from text. A keyframe is that still. It locks the face, the outfit and the framing, so the video model only has to add motion.
- Make a 9:16 keyframe for the shot from your character sheet. See character consistency.
- Load it into the Wan 2.2 I2V template.
- Set width 720 and height 1280 for vertical. The Wan FLF docs warn that the 720p models give poor results at small sizes.
- Set the frame count. Wan 2.2 A14B defaults to 81 frames at 16 fps on fal, which is about 5 seconds. One creator found Wan 2.2 misbehaves at 24 fps, so he switched his workflows back to 16.
- Prompt the motion and the camera, not the character's looks. The image already carries those.
Recipe 2: first-last frame for exact moves
Use this when a beat needs an exact ending: she sits down across from him, the door closes, he ends in a close-up. Make both stills, then run the FLF2V template.
It's also the cleanest way to chain shots. Save the last frame of one clip and use it as the first frame of the next. A Google Flow creator uses the same trick, and it works the same way here.
Recipe 3: a LoRA for each main character
A LoRA is a small add-on file you train on one character's face. When a lead shows up in 400 shots, it pays for itself.
- Train it with musubi-tuner (HunyuanVideo, Wan 2.1 and 2.2) or ai-toolkit (Wan 2.1 and 2.2, LTX-2.x, MiniMax H3).
- Load it with a Load LoRA node. The docs explain
strength_modelas how strongly the LoRA affects the model. Start around the middle and adjust. - Add it to both Wan 2.2 experts. Wan 2.2 uses two models, a high-noise one and a low-noise one, and a tutorial shows adding the LoRA to both. Missing one is a common reason a LoRA seems to do nothing.
Train on that one character only. Use varied lighting, angles and expressions, in the outfit they wear most.
Recipe 4: dialogue shots
You have two local options:
- Wan 2.2 S2V: give it a still and the line's audio. Each chunk is about 4.8 seconds, so match the number of extend nodes to your audio length. The docs show the math: a 14-second line needs 224 frames at 16 fps.
- InfiniteTalk: driven by audio and built for long talking clips with head and body motion. Expect it to need more VRAM than basic Wan workflows.
Make the voice first with TTS (text-to-speech) or an actor. Then make the shot. More in lip-sync and dubbing.
Recipe 5: act it out yourself
Wan 2.2 Animate's Move mode copies the motion and expressions from a video of you onto your character image. Film yourself playing the scene on a phone, vertical, against a plain wall. For a confrontation with tricky timing, this is often faster than prompting.
Finish: upscale and smooth
- Upscale (make the video sharper and bigger). The official upscale workflow uses Load Upscale Model and Upscale Image (using Model) nodes with ESRGAN-type models. On video it runs on every frame, so only do it on final takes.
- Frame interpolation (add in-between frames). ComfyUI-Frame-Interpolation (MIT) adds frames. It's handy for taking 16 fps Wan output to a smoother rate.
- Video in and out. VideoHelperSuite loads and saves video, and most community workflows use it.
LTX-2 also ships spatial and temporal 2x upscalers. They're listed in its ComfyUI docs.
Batch every shot in an episode
Don't click Queue 40 times per episode. ComfyUI's server takes jobs over HTTP, so you can send every shot from a script.
Export your workflow in API format, change the prompt and image for each shot, and post it to /prompt. Per the server routes docs, /history returns results and /ws streams progress. Here's a minimal loop, based on the pattern in ComfyUI's own docs:
import json, copy
from urllib import request
SERVER = "http://127.0.0.1:8188"
base = json.load(open("wan22_i2v_api.json")) # exported in API format
shots = [
{"id": "ep03_sc02_sh01", "image": "ep03_sc02_sh01.png",
"prompt": "slow push-in, she looks up from the contract, rain on the window"},
{"id": "ep03_sc02_sh02", "image": "ep03_sc02_sh02.png",
"prompt": "static close-up, he exhales and loosens his cufflink"},
]
for shot in shots:
wf = copy.deepcopy(base)
wf["6"]["inputs"]["text"] = shot["prompt"] # your positive prompt node id
wf["52"]["inputs"]["image"] = shot["image"] # your LoadImage node id
wf["9"]["inputs"]["filename_prefix"] = shot["id"] # your save node id
data = json.dumps({"prompt": wf}).encode("utf-8")
request.urlopen(request.Request(f"{SERVER}/prompt", data=data))
Node IDs differ per workflow, so open your exported JSON and find yours. Name outputs by episode, scene and shot. Then your editor can drop them straight onto the timeline.
Speed tricks creators use
- Step-distill LoRAs. These let the model finish in fewer steps. The Wan 2.2 Animate template ships with a lightx2v 4-step acceleration LoRA, and similar LoRAs exist for other Wan workflows.
- Quantized models. GGUF versions are compressed models. They lose a little quality but run on smaller cards.
- Sage Attention. It's a faster way for the model to do its math. In one tutorial it worked on RTX 30, 40 and 50 series cards but not on a 2060.
- Draft small, finish big. Run 480p drafts to settle the prompt and timing, then re-run only the winner at 720p. A new resolution won't copy the draft exactly, so plan a retake or two.
Keep it organized
Save one workflow per shot type: close-up I2V, walk-and-talk FLF, dialogue S2V. For every approved shot, log the seed (the number that sets the random start), the LoRA strength and the prompt. Keep that spreadsheet next to your shot list.
When episode 47 needs a pickup shot that matches episode 12, you'll be glad you did. More community workflows are under workflow projects, and the storyboarding guide covers planning shots before you open ComfyUI.
Sources
- github.com/Comfy-Org/ComfyUI
- docs.comfy.org/installation/system_requirements
- docs.comfy.org/tutorials/video/wan/wan2_2
- docs.comfy.org/tutorials/video/wan/wan-flf
- docs.comfy.org/tutorials/video/wan/wan2-2-s2v
- docs.comfy.org/tutorials/video/wan/wan2-2-animate
- docs.comfy.org/tutorials/video/ltx/ltx-2
- docs.comfy.org/tutorials/video/hunyuan/hunyuan-video-1-5
- docs.comfy.org/tutorials/basic/lora
- docs.comfy.org/tutorials/basic/upscale
- docs.comfy.org/development/comfyui-server/comms_routes
- docs.comfy.org/development/comfyui-server/api-key-integration
- github.com/Wan-Video/Wan2.2
- github.com/kijai/ComfyUI-WanVideoWrapper
- github.com/Fannovel16/ComfyUI-Frame-Interpolation
- github.com/Kosinkadink/ComfyUI-VideoHelperSuite
- github.com/kohya-ss/musubi-tuner
- github.com/ostris/ai-toolkit
- github.com/MeiGen-AI/InfiniteTalk
- fal.ai/models/fal-ai/wan/v2.2-a14b/text-to-video
- www.youtube.com/watch?v=Z8JlJdXdVg4
- www.youtube.com/watch?v=CmHj--QJ8U4