Skip to content

modelscope /DiffSynth-Studio

Trains and runs image, video and audio diffusion models on your own GPU
Model toolingApache-2.0English README
Stars
13k
30 days
New
Last push
yesterday

About DiffSynth-Studio

DiffSynth-Studio runs open-source image, video and audio models on your own GPU. The ModelScope community team built it, and it covers both running models and training them. Low-VRAM options and quantization let consumer cards handle big models. A small maintainer team means new features and issue replies arrive slowly.

Best for: Developers and researchers who want to run or fine-tune big generative models on one consumer GPU.

What it does

  • Runs many open-source image, video and audio diffusion models
  • Moves model weights between disk, memory and GPU to fit small VRAM
  • Quantizes weights to NF4 or INT8 to cut VRAM use
  • Trains base models, LoRAs and adapter models
  • Splits training into two stages to save memory
  • Pairs with DiffSynth-ComfyUI nodes and DiffSynth-WebUI for LoRA training

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Open the README and find the model you want in the supported models list.
  2. 2
    Go to that model's folder under examples/ and copy the example script.
  3. 3
    Run the script on your own GPU. If you run out of memory, use the low-VRAM or quantized example instead.

Models and languages

Models it works with

  • MiniMax H3
  • LTX-2.5
  • Wan-Animate-2
  • LingBot-Video
  • Qwen-Video-Edit
  • Qwen-Image-2.1
  • Z-Image-Turbo
  • FLUX.2-klein-base-4B
  • HiDream-O1-Image
  • ACE-Step

Interface and docs

  • Chinese
  • English

Head-to-head

Alternatives

Other model tooling projects people compare with DiffSynth-Studio.

ModelTC

LightX2V

Loads open image and video models onto your own GPU

Model tooling
  • MiniMax H3
  • Wan 2.2
  • Wan 2.1
  • +7
2.9k GitHub starstodayApache-2.0

aigc-apps

VideoX-Fun

Python toolkit that makes AI video and trains your own models

Model tooling
  • CogVideoX-Fun V1.1 2B
  • CogVideoX-Fun V1.1 5B
  • Wan 2.1-Fun 1.3B
  • +6
2.3k GitHub stars2 days agoApache-2.0

mrbizarro

Phosphene

Video, image and music models on your Mac, no cloud or API keys

Model tooling
  • LTX-Video 2.5
  • MiniMax H3
  • Qwen-Image-Edit-2509
  • +3
242 GitHub starstodayMIT

Wan-Video

Wan2.2

Open video models for text, image, speech and character animation

Model tooling
  • Wan 2.2-T2V-A14B
  • Wan 2.2-I2V-A14B
  • Wan 2.2-TI2V-5B
  • +5
18k GitHub stars10 days agoApache-2.0

Featured on OpenMicroDrama

Do you maintain DiffSynth-Studio? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/modelscope-diffsynth-studio)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Built by the ModelScope community team, which also runs hosted products powered by it. We write these descriptions ourselves. The repo's own docs are the final word.