RVC-Boss
GPT-SoVITS
Clones a voice from a 5-second clip and reads your text in it
- GPT-SoVITS v2Pro / v2ProPlus
- GPT-SoVITS v4
- GPT-SoVITS v3
- +7

MOSS-TTSD is an open text-to-speech model built for dialogue, not one-voice narration. Give it a short clip of each speaker, then a script tagged [S1] to [S5]. It keeps each voice steady for up to 60 minutes of audio. The documented setup needs a GPU with flash-attention.
Best for: Voice directors who dub multi-speaker scripts into podcasts, audiobooks or drama dialogue.
Rewritten from the README. Check the repo for the latest steps.
conda create -n moss_ttsd python=3.12 -y && conda activate moss_ttsdpip install flash-attnpip install -r requirements.txt[S1]–[S5] speaker tags and matching prompt_audio_speakerN / prompt_text_speakerN reference pairspython inference.py --model_path OpenMOSS-Team/MOSS-TTSD-v1.0 --codec_model_path OpenMOSS-Team/MOSS-Audio-Tokenizer --input_jsonl /path/to/input.jsonl --save_dir outputs --mode voice_clone_and_continuation --batch_size 1 --text_normalizepython scripts/request_sglang_generation.pyOther editing & audio projects people compare with MOSS-TTSD.
RVC-Boss
Clones a voice from a 5-second clip and reads your text in it
index-tts
Voice cloner that reads your script in five languages
hkchengrex
Makes sound effects and ambience that match your video
jianchang512
Video translated, subtitled and dubbed into another language
Do you maintain MOSS-TTSD? Add this badge to your README so English speakers can find our write-up.
[](https://openmicrodrama.com/projects/openmoss-moss-ttsd)The badge links to this page. Want your description changed? Email support@openmicrodrama.com.
Reviewed Oct 1, 2026. We write these descriptions ourselves. The repo's own docs are the final word.