Skip to content

OpenMOSS /MOSS-TTSD

Cloned-voice dialogue audio for 1 to 5 speakers
Editing & audioApache-2.0English README
Stars
1.4k
30 days
New
Last push
25 days ago

About MOSS-TTSD

MOSS-TTSD is an open text-to-speech model built for dialogue, not one-voice narration. Give it a short clip of each speaker, then a script tagged [S1] to [S5]. It keeps each voice steady for up to 60 minutes of audio. The documented setup needs a GPU with flash-attention.

Best for: Voice directors who dub multi-speaker scripts into podcasts, audiobooks or drama dialogue.

What it does

  • Clones 1 to 5 voices from short reference clips, no training needed
  • Keeps each speaker's voice steady for up to 60 minutes in one run
  • Speaks 20 languages, including Chinese, English and Japanese
  • Runs batch jobs from a JSONL file of tagged dialogue scripts
  • Serves the model through SGLang for faster generation
  • Ships fine-tuning code for full-parameter and LoRA training (v0.5)

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Create the environment: the command below
    conda create -n moss_ttsd python=3.12 -y && conda activate moss_ttsd
  2. 2
    Install dependencies: the command below then pip install flash-attn
    pip install -r requirements.txt
  3. 3
    Write a JSONL file with [S1]–[S5] speaker tags and matching prompt_audio_speakerN / prompt_text_speakerN reference pairs
  4. 4
    Run batch inference: the command below
    python inference.py --model_path OpenMOSS-Team/MOSS-TTSD-v1.0 --codec_model_path OpenMOSS-Team/MOSS-Audio-Tokenizer --input_jsonl /path/to/input.jsonl --save_dir outputs --mode voice_clone_and_continuation --batch_size 1 --text_normalize
  5. 5
    Optional: serve the fused model with SGLang and send requests via the command below
    python scripts/request_sglang_generation.py

Models and languages

Models it works with

  • MOSS-TTSD v1.0
  • MOSS-Audio-Tokenizer
  • XY-Tokenizer
  • MOSS-TTSD v0.7
  • MOSS-TTSD v0.5

Interface and docs

  • English
  • Chinese

Alternatives

Other editing & audio projects people compare with MOSS-TTSD.

RVC-Boss

GPT-SoVITS

Clones a voice from a 5-second clip and reads your text in it

Editing & audio
  • GPT-SoVITS v2Pro / v2ProPlus
  • GPT-SoVITS v4
  • GPT-SoVITS v3
  • +7
62k GitHub stars1 month agoMIT

index-tts

index-tts

Voice cloner that reads your script in five languages

Editing & audio
  • IndexTTS-2.5
  • IndexTTS-2
  • IndexTTS-1.5
  • +2
24k GitHub starsyesterdayCustom license

hkchengrex

MMAudio

Makes sound effects and ambience that match your video

Editing & audio
  • MMAudio large_44k_v2
2.3k GitHub stars7 months agoMIT

k4yt3x

video2x

Makes small video bigger and smoother using Anime4K, Real-ESRGAN, RIFE

Editing & audio
  • Anime4K v4
  • Real-ESRGAN
  • Real-CUGAN
  • +1
22k GitHub stars6 months agoAGPL-3.0

Featured on OpenMicroDrama

Do you maintain MOSS-TTSD? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/openmoss-moss-ttsd)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. We write these descriptions ourselves. The repo's own docs are the final word.