Skip to content

MeiGen-AI /InfiniteTalk

Lip-synced video or photo from new audio, any length
Editing & audioApache-2.0English README
Stars
8k
30 days
New
Last push
4 months ago

About InfiniteTalk

Give InfiniteTalk a video or still photo plus audio, and it makes a new video where mouth, head and body follow the sound. Streaming mode keeps going past one minute for long dubs, though you'll need a big GPU and some setup. After about a minute, colors drift and the face matches less well.

Best for: Dubbing editors who want to re-animate existing footage or a portrait photo with matching lip sync.

What it does

  • Syncs lips, head, body and expressions to your audio track
  • Dubs a video, or animates one photo with an audio track
  • Streams unlimited length; clip mode makes one short chunk
  • Exports 480P or 720P video
  • Speeds up with TeaCache, int8 quantization and FusionX or lightx2v LoRAs
  • Runs on several GPUs, or on very low VRAM with a flag

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Make a conda env named multitalk and install PyTorch 2.4.1 and xformers 0.0.28 with CUDA 12.1 wheels.
  2. 2
    Install flash-attn 2.7.4.post1, then run the command below and the next command.
    pip install -r requirements.txt
    conda install -c conda-forge librosa
  3. 3
    Install ffmpeg, then download the base model, audio encoder and InfiniteTalk weights with huggingface-cli download.
  4. 4
    Run single-GPU inference: the command below
    python generate_infinitetalk.py --ckpt_dir weights/Wan2.1-I2V-14B-480P --wav2vec_dir 'weights/chinese-wav2vec2-base' --infinitetalk_dir weights/InfiniteTalk/single/infinitetalk.safetensors --input_json examples/single_example_image.json --size infinitetalk-480 --sample_steps 40 --mode streaming --motion_frame 9 --save_file infinitetalk_res
  5. 5
    Or launch the Gradio demo: the command below
    python app.py --ckpt_dir weights/Wan2.1-I2V-14B-480P --wav2vec_dir 'weights/chinese-wav2vec2-base' --infinitetalk_dir weights/InfiniteTalk/single/infinitetalk.safetensors --num_persistent_param_in_dit 0 --motion_frame 9

Models and languages

Models it works with

  • Wan 2.1-I2V-14B-480P
  • chinese-wav2vec2-base
  • InfiniteTalk
  • FusionX LoRA
  • lightx2v LoRA
  • LongCat-Video-Avatar

Interface and docs

  • English

Alternatives

Other editing & audio projects people compare with InfiniteTalk.

KlingAIResearch

X-Dub

Tool that redraws a video character's mouth to match a new audio track

Editing & audio
  • Wan 2.2 TI2V-5B
  • Whisper large-v2
  • Wav2Vec2-base-960h
  • +2
235 GitHub stars1 month agoApache-2.0

bytedance

LatentSync

Re-syncs a talking-face video to another audio track

Editing & audio
  • Stable Diffusion
  • Whisper tiny
  • SyncNet
6.1k GitHub stars1 year agoApache-2.0

KlingAIResearch

LivePortrait

Animates a face photo or video with motion from another clip

Editing & audio
19k GitHub stars4 months agoMIT

RVC-Boss

GPT-SoVITS

Clones a voice from a 5-second clip and reads your text in it

Editing & audio
  • GPT-SoVITS v2Pro / v2ProPlus
  • GPT-SoVITS v4
  • GPT-SoVITS v3
  • +7
62k GitHub stars1 month agoMIT

Featured on OpenMicroDrama

Do you maintain InfiniteTalk? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/meigen-ai-infinitetalk)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Last push 2026-05-22, outside our 90-day window. We kept it because it's still widely used. From the MeiGen-AI team; the same team has since released LongCat-Video-Avatar-1.5, an upgraded model. We write these descriptions ourselves. The repo's own docs are the final word.