Skip to content

WeChatCV /Stand-In

Wan add-on that keeps one face consistent across shots
Model toolingApache-2.0English README

About Stand-In

Stand-In adds a consistent face to Wan video generation from one photo. It trains only about 1% extra parameters, so the base model stays frozen. You get scripts for text-to-video, LoRA styles, face swapping and VACE pose control. Training code and dataset aren't out yet, and face swapping is experimental.

Best for: Technical artists who need one face to stay consistent across Wan-generated shots.

What it does

  • Keeps a character's face the same across shots from one reference photo
  • Trains about 1% extra parameters and leaves the base model frozen
  • Loads community LoRA (small add-on style models) for stylized video
  • Swaps faces in existing video with a denoising-strength setting
  • Works with VACE (a control model) for pose or depth control plus identity
  • Downloads base, face-recognition and Stand-In weights in one script

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Clone the repo and make a Python 3.11 environment: the command below, the next command, conda activate Stand-In
    git clone https://github.com/WeChatCV/Stand-In.git
    conda create -n Stand-In python=3.11 -y
  2. 2
    Install the requirements: the command below
    pip install -r requirements.txt
  3. 3
    Download the weights into checkpoints: python download_models.py (add --wan_version 2.2 for Wan2.2, --vace for VACE)
  4. 4
    Generate a video: the command below
    python infer.py --prompt "..." --ip_image "test/input/lecun.jpg" --output "test/output/lecun.mp4"
  5. 5
    For a style, run python infer_with_lora.py with --lora_path and --lora_scale
  6. 6
    For face swapping, run python infer_face_swap.py with --denoising_strength

Models and languages

Models it works with

  • Wan 2.1 T2V 14B
  • Wan 2.2 T2V A14B
  • VACE
  • AntelopeV2
  • Stand-In v1.0

Interface and docs

  • English

Alternatives

Other model tooling projects people compare with Stand-In.

Kevin-thu

StoryMem

Give it a shot list, get a minute-long multi-shot video

Model tooling
  • Wan 2.2 T2V-A14B
  • Wan 2.2 I2V-A14B
  • StoryMem Wan2.2 M2V-A14B
  • +1
- GitHub starsCustom license

KlingAIResearch

ShotStream

Chains multi-shot video frames on the fly for interactive stories

Model tooling
  • Wan 2.1-T2V-1.3B
- GitHub starsNo license

aigc-apps

VideoX-Fun

Python toolkit that makes AI video and trains your own models

Model tooling
  • CogVideoX-Fun V1.1 2B
  • CogVideoX-Fun V1.1 5B
  • Wan 2.1-Fun 1.3B
  • +6
- GitHub starsApache-2.0

Wan-Video

Wan2.2

Open video models for text, image, speech and character animation

Model tooling
  • Wan 2.2-T2V-A14B
  • Wan 2.2-I2V-A14B
  • Wan 2.2-TI2V-5B
  • +5
- GitHub starsApache-2.0

Featured on OpenMicroDrama

Do you maintain Stand-In? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/wechatcv-stand-in)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Built on Wan2.1 and DiffSynth-Studio. The third-party ComfyUI node differs from the official version. We write these descriptions ourselves. The repo's own docs are the final word.