Skip to content

bytedance /LatentSync

Re-syncs a talking-face video to another audio track
Editing & audioApache-2.0English README

About LatentSync

LatentSync re-syncs a talking-face video to a new audio track. ByteDance's open-source lip-sync model uses diffusion (the method behind Stable Diffusion). It ships with a Gradio app, a command-line script, training code and eval tools. Inference needs 8-18 GB of VRAM, and you download the checkpoints yourself.

Best for: Researchers and tinkerers who want to dub talking-head or anime clips with a diffusion lip-sync model.

What it does

  • Re-syncs lips to new audio without a separate motion step
  • Turns audio into embeddings with Whisper, then feeds them to the model
  • Ships a Gradio web app and a command-line inference script
  • Cleans training data: scene cuts, face alignment, quality filters
  • Trains the model and SyncNet with configs for different VRAM sizes
  • Checks sync confidence and SyncNet accuracy on your clips

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Run source setup_env.sh to install packages and download the checkpoints.
  2. 2
    Start the Gradio demo with python gradio_app.py.
  3. 3
    Or run command-line inference with ./inference.sh.
  4. 4
    Raise inference_steps (20-50) for sharper video, or guidance_scale (1.0-3.0) for tighter lip-sync.
  5. 5
    To train, process your data with ./data_processing_pipeline.sh, then run ./train_unet.sh or ./train_syncnet.sh.

Models and languages

Models it works with

  • Stable Diffusion
  • Whisper tiny
  • SyncNet

Interface and docs

  • English

Alternatives

Other editing & audio projects people compare with LatentSync.

KlingAIResearch

LivePortrait

Animates a face photo or video with motion from another clip

Editing & audio
- GitHub starsMIT

MeiGen-AI

InfiniteTalk

Lip-synced video or photo from new audio, any length

Editing & audio
  • Wan 2.1-I2V-14B-480P
  • chinese-wav2vec2-base
  • InfiniteTalk
  • +3
- GitHub starsApache-2.0

KlingAIResearch

X-Dub

Tool that redraws a video character's mouth to match a new audio track

Editing & audio
  • Wan 2.2 TI2V-5B
  • Whisper large-v2
  • Wav2Vec2-base-960h
  • +2
- GitHub starsApache-2.0

Huanshere

VideoLingo

Translated subtitles and a dubbed voice track for any video

Editing & audio
  • Qwen3-ASR
  • Qwen3-ForcedAligner
  • ElevenLabs
  • +7
- GitHub starsApache-2.0

Featured on OpenMicroDrama

Do you maintain LatentSync? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/bytedance-latentsync)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Last push 2025-06-20, outside our 90-day window. We kept it because it's still widely used. Built on AnimateDiff; borrows code from MuseTalk, StyleSync, SyncNet and Wav2Lip. We write these descriptions ourselves. The repo's own docs are the final word.