Skip to content

RVC-Boss /GPT-SoVITS

Clones a voice from a 5-second clip and reads your text in it
Editing & audioMITEnglish README
Stars
62k
30 days
New
Last push
1 month ago

About GPT-SoVITS

Give GPT-SoVITS a 5-second clip and it reads your text in that voice right away. About a minute of audio gets you a closer match. On a Mac, training runs on CPU for now, since Mac GPU training gave poor results.

Best for: Dubbing teams building character voices from short reference clips.

What it does

  • Clones a voice from a 5-second sample, no training needed
  • Trains on about 1 minute of audio for a closer voice match
  • Speaks English, Japanese, Korean, Cantonese and Chinese
  • Splits vocals from music, slices audio and writes transcripts
  • Picks from Fun-ASR-Nano, SenseVoice or classic FunASR for transcription
  • Runs in a WebUI or through API scripts

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Make a conda environment: the command below then conda activate GPTSoVits
    conda create -n GPTSoVits python=3.10
  2. 2
    On Linux or macOS run the command below
    bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope> [--download-uvr5]
  3. 3
    On Windows run the command below
    pwsh -F install.ps1 --Device <CU126|CU128|CPU> --Source <HF|HF-Mirror|ModelScope> [--DownloadUVR5]
  4. 4
    If install.sh worked, skip the model download; otherwise put pretrained models in GPT_SoVITS/pretrained_models
  5. 5
    Start the WebUI with python webui.py, or double-click go-webui.bat on Windows
  6. 6
    For inference only, run the command below
    python GPT_SoVITS/inference_webui.py

Models and languages

Models it works with

  • GPT-SoVITS v2Pro / v2ProPlus
  • GPT-SoVITS v4
  • GPT-SoVITS v3
  • GPT-SoVITS v2
  • GPT-SoVITS v1
  • Fun-ASR-Nano
  • SenseVoice
  • FunASR
  • Faster Whisper Large V3
  • UVR5

Interface and docs

  • English
  • Chinese
  • Japanese
  • Korean
  • Turkish

Head-to-head

Alternatives

Other editing & audio projects people compare with GPT-SoVITS.

OpenMOSS

MOSS-TTSD

Cloned-voice dialogue audio for 1 to 5 speakers

Editing & audio
  • MOSS-TTSD v1.0
  • MOSS-Audio-Tokenizer
  • XY-Tokenizer
  • +2
1.4k GitHub stars25 days agoApache-2.0

index-tts

index-tts

Voice cloner that reads your script in five languages

Editing & audio
  • IndexTTS-2.5
  • IndexTTS-2
  • IndexTTS-1.5
  • +2
24k GitHub starsyesterdayCustom license

hkchengrex

MMAudio

Makes sound effects and ambience that match your video

Editing & audio
  • MMAudio large_44k_v2
2.3k GitHub stars7 months agoMIT

k4yt3x

video2x

Makes small video bigger and smoother using Anime4K, Real-ESRGAN, RIFE

Editing & audio
  • Anime4K v4
  • Real-ESRGAN
  • Real-CUGAN
  • +1
22k GitHub stars6 months agoAGPL-3.0

Featured on OpenMicroDrama

Do you maintain GPT-SoVITS? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/rvc-boss-gpt-sovits)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. We write these descriptions ourselves. The repo's own docs are the final word.