Skip to content

KlingAIResearch /X-Dub

Tool that redraws a video character's mouth to match a new audio track
Editing & audioApache-2.0English README
Stars
235
30 days
New
Last push
1 month ago

About X-Dub

X-Dub takes a video plus an audio file and redraws the mouth to match. It's the official code for a research paper, and the public build runs on Wan2.2-TI2V-5B. Budget about 21 GB of VRAM, and it handles one person per clip. The public release is weaker than the paper version: flicker, drift (the face slowly changing) and noisy frames.

Best for: Re-dub artists who need to match one on-screen character's mouth to a new audio track.

What it does

  • Dubs one on-screen person's mouth to a new audio track
  • Skips masking: no manual mouth-region mask needed
  • Auto-crops any video size, then pastes the result back
  • Uses DWPose (a pose tool that finds face points) to place the crop
  • Lets you tune ref_cfg_scale, audio_cfg_scale and 25-50 steps
  • Runs inside ComfyUI through community-made nodes

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Clone the repo and make a Python 3.10 env: the command below, the next command, conda activate x-dub
    git clone https://github.com/KlingAIResearch/X-Dub.git
    conda create -n x-dub python=3.10 -y
  2. 2
    Install dependencies: the command below, then the OpenMMLab packages (mmengine, mmcv, mmdet, mmpose, chumpy), then pip install -e . --no-deps
    pip install -r requirements.txt
  3. 3
    Download the weights into checkpoints/: the command below
    hf download KlingTeam/X-Dub --local-dir ./checkpoints --repo-type model
  4. 4
    Move the DWPose files: mkdir -p dwpose_tools/models then the command below
    cp -r ./checkpoints/dwpose_tools/models/. ./dwpose_tools/models/
  5. 5
    Run inference: the command below
    python infer_lip_sync_pipeline.py --video_path assets/examples/video.mp4 --audio_path assets/examples/audio.wav --ckpt_path checkpoints/X-Dub_model.safetensors --ref_cfg_scale 2.5 --audio_cfg_scale 10.0 --num_inference_steps 30 --output_dir ./results

Models and languages

Models it works with

  • Wan 2.2 TI2V-5B
  • Whisper large-v2
  • Wav2Vec2-base-960h
  • UMT5-XXL
  • DWPose

Interface and docs

  • English

Alternatives

Other editing & audio projects people compare with X-Dub.

MeiGen-AI

InfiniteTalk

Lip-synced video or photo from new audio, any length

Editing & audio
  • Wan 2.1-I2V-14B-480P
  • chinese-wav2vec2-base
  • InfiniteTalk
  • +3
8k GitHub stars4 months agoApache-2.0

KlingAIResearch

LivePortrait

Animates a face photo or video with motion from another clip

Editing & audio
19k GitHub stars4 months agoMIT

bytedance

LatentSync

Re-syncs a talking-face video to another audio track

Editing & audio
  • Stable Diffusion
  • Whisper tiny
  • SyncNet
6.1k GitHub stars1 year agoApache-2.0

RVC-Boss

GPT-SoVITS

Clones a voice from a 5-second clip and reads your text in it

Editing & audio
  • GPT-SoVITS v2Pro / v2ProPlus
  • GPT-SoVITS v4
  • GPT-SoVITS v3
  • +7
62k GitHub stars1 month agoMIT

Featured on OpenMicroDrama

Do you maintain X-Dub? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/klingairesearch-x-dub)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. The paper's internal model isn't open source; this public build runs on Wan2.2-TI2V-5B and is adapted from DiffSynth-Studio. We write these descriptions ourselves. The repo's own docs are the final word.