OpenMOSS
MOSS-TTSD
Cloned-voice dialogue audio for 1 to 5 speakers
- MOSS-TTSD v1.0
- MOSS-Audio-Tokenizer
- XY-Tokenizer
- +2

MMAudio makes sound for a video clip, or from a text prompt alone. It lines the audio up with what's happening on screen. You'll need a GPU with about 6 GB of memory. The pretrained models are non-commercial, so check before you sell anything.
Best for: Sound designers who want automatic sound effects or ambience for short clips.
Rewritten from the README. Check the repo for the latest steps.
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118 --upgradegit clone https://github.com/hkchengrex/MMAudio.gitcd MMAudio then pip install -e .--video for text-to-audio.python demo.py --duration=8 --video= --prompt "your prompt"python gradio_demo.py (default port 7860, set with --port)Other editing & audio projects people compare with MMAudio.
OpenMOSS
Cloned-voice dialogue audio for 1 to 5 speakers
RVC-Boss
Clones a voice from a 5-second clip and reads your text in it
index-tts
Voice cloner that reads your script in five languages
jianchang512
Video translated, subtitled and dubbed into another language
Do you maintain MMAudio? Add this badge to your README so English speakers can find our write-up.
[](https://openmicrodrama.com/projects/hkchengrex-mmaudio)The badge links to this page. Want your description changed? Email support@openmicrodrama.com.
Reviewed Oct 1, 2026. Last push 2026-02-23, outside our 90-day window. We kept it because it's still widely used. Not affiliated with the site mmaudio.net. We write these descriptions ourselves. The repo's own docs are the final word.