RVC-Boss
GPT-SoVITS
Clones a voice from a 5-second clip and reads your text in it
- GPT-SoVITS v2Pro / v2ProPlus
- GPT-SoVITS v4
- GPT-SoVITS v3
- +7
MOSS-TTSD is an open text-to-speech model built for dialogue, not one-voice narration. Give it a short clip of each speaker, then a script tagged [S1] to [S5]. It keeps each voice steady for up to 60 minutes of audio. The documented setup needs a GPU with flash-attention.
Best for: Voice directors who dub multi-speaker scripts into podcasts, audiobooks or drama dialogue.
Rewritten from the README. Check the repo for the latest steps.
conda create -n moss_ttsd python=3.12 -y && conda activate moss_ttsdpip install flash-attnpip install -r requirements.txt[S1]–[S5] speaker tags and matching prompt_audio_speakerN / prompt_text_speakerN reference pairspython inference.py --model_path OpenMOSS-Team/MOSS-TTSD-v1.0 --codec_model_path OpenMOSS-Team/MOSS-Audio-Tokenizer --input_jsonl /path/to/input.jsonl --save_dir outputs --mode voice_clone_and_continuation --batch_size 1 --text_normalizepython scripts/request_sglang_generation.pyOther editing & audio projects people compare with MOSS-TTSD.
RVC-Boss
Clones a voice from a 5-second clip and reads your text in it
index-tts
Voice cloner that reads your script in five languages
hkchengrex
Makes sound effects and ambience that match your video
k4yt3x
Makes small video bigger and smoother using Anime4K, Real-ESRGAN, RIFE
Do you maintain MOSS-TTSD? Add this badge to your README so English speakers can find our write-up.
[](https://openmicrodrama.com/projects/openmoss-moss-ttsd)The badge links to this page. Want your description changed? Email support@openmicrodrama.com.
Reviewed Oct 1, 2026. We write these descriptions ourselves. The repo's own docs are the final word.