Skip to content

aigc-apps /VideoX-Fun

Python toolkit that makes AI video and trains your own models
Model toolingApache-2.0English README
Stars
2.3k
30 days
New
Last push
2 days ago

About VideoX-Fun

VideoX-Fun makes AI images and videos, and trains your own video models. It runs text-to-video, image-to-video, video-to-video, and control videos from Canny, Pose or Depth. You drive it with Python scripts, a Gradio web UI, or ComfyUI nodes. Weights take about 60GB of disk, and big models need offload or many GPUs.

Best for: Builders and researchers who want to run or fine-tune open video models on their own GPUs.

What it does

  • Runs text-to-video, image-to-video, video-to-video and control videos
  • Uses one script entry for video and image models under examples/{model_name}/
  • Saves GPU memory with model_cpu_offload, qfloat8 or sequential_cpu_offload
  • Splits work across GPUs with xfuser (ulysses_degree and ring_degree)
  • Trains baseline, LoRA (a small add-on model), control and distill models
  • Prepares data: splits long videos, cleans them and writes captions

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Pull the Docker image the command below and run it with --gpus all.
    docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
  2. 2
    Clone the repo with the command below, then cd VideoX-Fun.
    git clone https://github.com/aigc-apps/VideoX-Fun.git
  3. 3
    Make models/Diffusion_Transformer and models/Personalized_Model, then download weights from Hugging Face or ModelScope into them.
  4. 4
    Edit the prompt and seed in examples/cogvideox_fun/predict_t2v.py, run it, and find videos in samples/cogvideox-fun-videos.
  5. 5
    Or run examples/cogvideox_fun/app.py to open the Gradio interface and generate there.
  6. 6
    For multi-GPU, install xfuser==0.4.2 and yunchang==0.6.2, then run the command below.
    torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py

Models and languages

Models it works with

  • CogVideoX-Fun V1.1 2B
  • CogVideoX-Fun V1.1 5B
  • Wan 2.1-Fun 1.3B
  • Wan 2.1-Fun 14B
  • Wan 2.2
  • Wan 2.2-Fun
  • Qwen-Image
  • Qwen-Image-2.1
  • Z-Image

Interface and docs

  • English
  • Chinese
  • Japanese

Alternatives

Other model tooling projects people compare with VideoX-Fun.

Trains and runs image, video and audio diffusion models on your own GPU

Model tooling
  • MiniMax H3
  • LTX-2.5
  • Wan-Animate-2
  • +7
13k GitHub starsyesterdayApache-2.0

ModelTC

LightX2V

Loads open image and video models onto your own GPU

Model tooling
  • MiniMax H3
  • Wan 2.2
  • Wan 2.1
  • +7
2.9k GitHub starstodayApache-2.0

Lightricks

LTX-Desktop

Local LTX video generation and editing on your desktop

Model tooling
  • LTX 2.5 Fast
  • LTX 2.5 Pro
  • LTX 2.3 Fast
  • +3
2k GitHub stars1 month agoApache-2.0

Wan-Video

Wan2.2

Open video models for text, image, speech and character animation

Model tooling
  • Wan 2.2-T2V-A14B
  • Wan 2.2-I2V-A14B
  • Wan 2.2-TI2V-5B
  • +5
18k GitHub stars10 days agoApache-2.0

Featured on OpenMicroDrama

Do you maintain VideoX-Fun? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/aigc-apps-videox-fun)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. We write these descriptions ourselves. The repo's own docs are the final word.