Build a character pack before you generate a single clip. Then give the model the same face images every time. That's what keeps your lead looking like the same person in episode 1 and episode 80.
Why it matters so much: the 9:16 frame is mostly close-ups, and people binge 20 episodes in one sitting. If your lead's jaw changes between episode 14 and 15, the comments will notice.
Below: what to make first, how each model takes reference images, and how to catch a changed face before your viewers do.
Why faces drift
Drift means the face slowly changing from shot to shot. It happens because every clip is a fresh guess by the model. Text alone doesn't pin a face down.
"Chestnut hair, cream blouse" fits thousands of people, so the model picks a new one each time.
ByteDance's own Seedance 2.0 guide names two common causes. One is a single reference image that crams face, pose, outfit and details together. The other is a face that's too small in the reference frame.
Their fix is simple. Add a separate head-only close-up next to the full-body image. Keep it expressionless, with very little shoulder and background.
Creators who make long AI videos say the same. Isa (Isa does AI) shows this in her long-form workflow video.
A saved character tag locks the face, but not the outfit, hair or styling. So she builds a full-body styled image as "the visual contract for the whole project."
Build a character pack first
A character pack is a small folder of images and audio for one character. Make one per character, per look, and name the folder after the character.
- Head-only close-up. Facing the camera, neutral face, even light, plain background. This is the image that holds the face in place best.
- Full-body styled image. The exact outfit, hair and accessories for this look. Make it in the same aspect ratio you'll use for video, so the model doesn't have to reinterpret it.
- Extra angles (optional). Save a three-quarter view and a profile, each as its own image. On Seedance 2.0, don't put several views on one sheet, because its guide warns it can cause "twins" in the frame.
- Seedance 2.5 is more flexible. Its guide says single-view and multi-view images both work for one to five people. Past five, it still says to use separate single-view images.
- Location plate. One clean image of each set you'll reuse, like the office, the mansion staircase or the pack lodge.
- Voice sample. A short, clean clip of the voice for this character. Only needed if your model or dubbing tool takes audio references.
Lock the descriptor string
Write one sentence that describes each character. Paste it word for word into every prompt, and never reword it. Here's the pair we use in the CEO romance prompts:
ADRIAN VALE: 32-year-old man, sharp jaw, slicked-back black hair, charcoal three-piece suit,
silver cufflinks, cold dark eyes.
ELENA BROOKS: 24-year-old woman, chestnut hair in a low ponytail, cream silk blouse,
thrifted tan trench coat.
The sentence does two jobs. It fills in what the image can't show, like eye color or a scar on the left side. It also gives you something to search for when you check 300 prompts later.
Keep a wardrobe bible
A wardrobe bible is a table of every outfit and when it's worn. Micro dramas change clothes on big story beats, like the makeover, the gala or the courtroom. The table stops episode 47 from bringing back the episode 3 coat.
| Character | Look | Episodes | Key pieces | Reference files |
|---|---|---|---|---|
| Elena | Day one | 1–9 | cream blouse, tan trench | elena-head.png, elena-look-a.png |
| Elena | Gala | 10–12 | emerald satin gown, hair down | elena-look-b.png |
| Adrian | Default | 1–80 | charcoal three-piece, cufflinks | adrian-head.png, adrian-look-a.png |
The takeaway: every look gets an episode range and its own reference file.
Reference inputs by model
A reference image is a picture you give the model so it copies that face or outfit. Every major model now takes some kind of character reference. The limits below come from each provider's API docs as of October 2026, and they change often.
| Model | How you reference a character | Limits worth knowing |
|---|---|---|
| Seedance 2.5 | Images, videos and audio, called in the prompt as @Image1, @Video1, @Audio1 | Up to 30 images, 10 videos, 10 audio clips. Official guide: 1 to 8 image subjects work best |
| Seedance 2.0 | Same tagging | Up to 9 images, 3 videos, 3 audio. Guide recommends 4 to 5 assets total, not the max |
| Kling 3.0 | "Elements": a frontal image plus 1 to 3 extra angles, or a video. A voice ID can be bound to an element | Start and end frames also supported, clips of 3 to 15 seconds |
| Veo 3.1 | Up to 3 reference images of a person, character or product | Clips are 4, 6 or 8 seconds; Veo videos can be extended 7 seconds at a time, up to 20 times |
| Hailuo (MiniMax H3) | Reference-to-video with images, video clips and audio, cited as Image 1, Video 1 | Up to 9 images and 3 videos for motion, 12 files total. LoRA training endpoints exist on fal |
| Wan 3.0 / 2.6 | Reference-to-video. Wan 2.6 R2V handles up to 3 referenced characters | Wan 3.0 takes up to 10 images, 5 videos, 5 audio. Wan 2.2 is open weights, so you can train a LoRA |
The takeaway: Seedance, MiniMax H3 and Wan 3.0 take the most references, and Veo takes the fewest. Check the linked pages before you plan a season around these numbers.
If you self-host, Wan 2.2 is the practical way to train a character LoRA. A LoRA is a small add-on you train on your character's face.
Wan 2.2 is open weights, which means you can download the model and run it yourself.
fal's hosted Wan 2.2 image-to-video trainer lists $5 for 1,000 steps. For local setups, see the Wan guide and ComfyUI for micro drama.
Write references into the prompt properly
Uploading a reference isn't enough. Tell the model what each file is for, and what it isn't for.
- Tie every file to a name. ByteDance's guide suggests putting the reference right after each character's name, the same way every time. For example: "Elena (Image 1) slides the folder toward Adrian (Image 2)."
- Say what to ignore. In a Seedance 2.5 tutorial, CottonByte AI notes that 2.5 copies references more strongly than 2.0. So say what each image should and shouldn't be used for, or its background can leak into your shot.
- Put the most important files first. The Seedance 2.0 guide says assets that need exact copying should come earlier in the prompt.
- Add a twin guard to group shots. The same guide suggests ending the prompt with a line that bans duplicate copies of a character in one frame.
- Name the style. If your references look real and you want anime, say so, or convert the references to anime first. The guide calls this problem "style drifting."
Keep shots connected
Chain frames. Take the last frame of one shot and use it as the first frame of the next. Kling, Veo and Wan all take a start image, and some take an end image too.
# grab the final frame of a clip as a PNG
ffmpeg -sseof -0.1 -i shot-07.mp4 -frames:v 1 shot-07-last.png
Use the last clip as a video reference. In Isa's long-form method, clip one goes into clip two as a reference, along with the styled character image. That carries the room and the motion forward.
Trim the joins. When you stitch or extend Seedance clips, ByteDance suggests cutting 6 frames from the end of the first clip. Then cut 1 frame from the start of the next one, so the jump is less visible.
Don't count on seeds to keep a face. A seed is a number that repeats the model's random starting point. But fal's Seedance schema notes results "may still vary slightly even with the same seed."
Use seeds to test one prompt change at a time, not to hold a face.
Don't switch model versions mid-season. Seedance 2.5's guide warns that 2.0 and 2.5 give clearly different looks from the same prompt and reference. Pick one version per series, or at least per story arc.
Voices are part of the character
Viewers notice a new voice faster than a small change in a face. So keep one voice per character for the whole season.
Kling lets you attach a voice to an element, and Seedance takes audio references. If you dub later, save the voice settings in the character pack. The lip-sync and dubbing guide has the tools.
Lip sync and face swap can change faces too
Every step after generation can nudge a face. Check the face again after lip sync (matching mouth movements to the words), after face repair and after upscaling.
Real-looking faces also hit platform rules. On BytePlus ModelArk, Seedance 2.0 and 2.5 don't accept uploaded references that show real human faces.
BytePlus lists three ways around it:
- face images its own models made under the same account within 30 days, unedited
- a preset library of digital characters
- real-person assets you're authorized to use
So a photoreal face sheet from another tool may get rejected. Test your pack on the platform you'll actually use, before you plan a season around it.
ByteDance's guide adds one more risk. A face that drifts can start to look like a celebrity and get blocked in review, so keep your references strong.
The copyright basics guide has more on likeness, meaning the right to your own face and voice.
Episode check before export
Run this list on every episode:
- The face matches the head-only reference in the first and last shot.
- Hair part, length and color match the wardrobe bible for these episodes.
- The outfit matches the look listed for this episode. Watch ties, earrings and coat buttons.
- No twins or extra copies of a character in group shots.
- Scars, tattoos and moles are on the correct side. Mirrored shots flip them.
- The voice sounds the same as in the last episode.
- Eye lines and screen direction hold across cuts (see vertical framing).
- No stray subtitles, logos or watermarks baked into the frame.
Where to go next
Plan your shots before you spend credits, with storyboarding with AI. Then grab prompts that already lock the cast for Kling, Seedance or Wan. Open-source character and storyboard tools are in our project directory.
Sources
- docs.byteplus.com/en/docs/ModelArk/2222480
- docs.byteplus.com/en/docs/ModelArk/2607689
- docs.byteplus.com/en/docs/ModelArk/2608626
- fal.ai/models/bytedance/seedance-2.5/reference-to-video
- fal.ai/models/bytedance/seedance-2.0/reference-to-video
- fal.ai/models/fal-ai/kling-video/v3/pro/image-to-video
- fal.ai/models/fal-ai/kling-video/o3/pro/reference-to-video
- ai.google.dev/gemini-api/docs/video
- fal.ai/models/minimax/h3/reference-to-video
- www.minimax.io/blog/minimax-h3
- fal.ai/models/alibaba/wan-3.0/reference-to-video
- www.alibabacloud.com/help/en/model-studio/video-generation
- fal.ai/models/fal-ai/wan-22-trainer/i2v-a14b
- www.youtube.com/watch?v=dOmKYJoRboE
- www.youtube.com/watch?v=vIksC7naeHo
- www.youtube.com/watch?v=96q3KHK9KEE
- www.youtube.com/watch?v=r9Z4GC7nb_8