Skip to content

hkchengrex /MMAudio

Makes sound effects and ambience that match your video
Editing & audioMITEnglish README
Stars
2.3k
30 days
New
Last push
7 months ago

About MMAudio

MMAudio makes sound for a video clip, or from a text prompt alone. It lines the audio up with what's happening on screen. You'll need a GPU with about 6 GB of memory. The pretrained models are non-commercial, so check before you sell anything.

Best for: Sound designers who want automatic sound effects or ambience for short clips.

What it does

  • Makes sound from a video, or from text alone when you skip the video
  • Lines up the generated audio with the video frames
  • Saves .flac audio and .mp4 video into ./output
  • Runs a Gradio web UI for video, text and image input
  • Fits in about 6 GB of GPU memory in 16-bit mode
  • Ships training, batch evaluation and onset evaluation scripts

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Install PyTorch 2.5.1+ with your CUDA build: the command below
    pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118 --upgrade
  2. 2
    Clone the repo: the command below
    git clone https://github.com/hkchengrex/MMAudio.git
  3. 3
    Install the package after PyTorch: cd MMAudio then pip install -e .
  4. 4
    Run the demo: the command below. Leave out --video for text-to-audio.
    python demo.py --duration=8 --video= --prompt "your prompt"
  5. 5
    Or start the Gradio app: python gradio_demo.py (default port 7860, set with --port)

Models and languages

Models it works with

  • MMAudio large_44k_v2

Interface and docs

  • English

Alternatives

Other editing & audio projects people compare with MMAudio.

OpenMOSS

MOSS-TTSD

Cloned-voice dialogue audio for 1 to 5 speakers

Editing & audio
  • MOSS-TTSD v1.0
  • MOSS-Audio-Tokenizer
  • XY-Tokenizer
  • +2
1.4k GitHub stars25 days agoApache-2.0

RVC-Boss

GPT-SoVITS

Clones a voice from a 5-second clip and reads your text in it

Editing & audio
  • GPT-SoVITS v2Pro / v2ProPlus
  • GPT-SoVITS v4
  • GPT-SoVITS v3
  • +7
62k GitHub stars1 month agoMIT

index-tts

index-tts

Voice cloner that reads your script in five languages

Editing & audio
  • IndexTTS-2.5
  • IndexTTS-2
  • IndexTTS-1.5
  • +2
24k GitHub starsyesterdayCustom license

k4yt3x

video2x

Makes small video bigger and smoother using Anime4K, Real-ESRGAN, RIFE

Editing & audio
  • Anime4K v4
  • Real-ESRGAN
  • Real-CUGAN
  • +1
22k GitHub stars6 months agoAGPL-3.0

Featured on OpenMicroDrama

Do you maintain MMAudio? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/hkchengrex-mmaudio)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Last push 2026-02-23, outside our 90-day window. We kept it because it's still widely used. Not affiliated with the site mmaudio.net. We write these descriptions ourselves. The repo's own docs are the final word.