Skip to content

KlingAIResearch /ShotStream

Chains multi-shot video frames on the fly for interactive stories
Model toolingNo licenseEnglish README
Stars
185
30 days
New
Last push
14 days ago

About ShotStream

ShotStream is research code that generates multi-shot video frames as you go. It's built on Wan2.1-T2V-1.3B and reports 16 FPS on one NVIDIA GPU. You get inference and training scripts, but it's a reference implementation, not an app. Training needs multi-node GPUs and edits to the bash scripts.

Best for: Researchers and engineers who want to run, fine-tune or study a streaming multi-shot video model.

What it does

  • Generates frames on the fly for interactive storytelling
  • Runs autoregressive 4-step long multi-shot inference
  • Reports 16 FPS on a single NVIDIA GPU
  • Trains a bidirectional next-shot teacher model
  • Initializes the causal student from teacher ODE pairs, following CausVid
  • Distills in two stages: intra-shot, then inter-shot self-forcing

Quickstart

Rewritten from the README. Check the repo for the latest steps.

  1. 1
    Clone the repo and make a conda environment: the command below, the next command, conda activate shotstream
    git clone https://github.com/KlingAIResearch/ShotStream.git
    conda create -n shotstream python=3.10 -y
  2. 2
    Install CUDA, PyTorch and the rest: the command below, then the next command. Or run bash tools/setup/env.sh
    pip install -r requirements.txt
    pip install flash-attn --no-build-isolation
  3. 3
    Download Wan-T2V-1.3B and ShotStream checkpoints into wan_models and ckpts with git-lfs. Or run the command below
    bash tools/setup/download_ckpt.sh
  4. 4
    Run 4-step long multi-shot generation: the command below
    bash tools/inference/causal_fewsteps.sh
  5. 5
    To train, set MASTER_ADDR in the bash scripts, then start with the command below
    bash tools/train/1_basemodel.sh 0

Models and languages

Models it works with

  • Wan 2.1-T2V-1.3B

Interface and docs

  • English

Alternatives

Other model tooling projects people compare with ShotStream.

WeChatCV

Stand-In

Wan add-on that keeps one face consistent across shots

Model tooling
  • Wan 2.1 T2V 14B
  • Wan 2.2 T2V A14B
  • VACE
  • +2
791 GitHub stars1 month agoApache-2.0

Kevin-thu

StoryMem

Give it a shot list, get a minute-long multi-shot video

Model tooling
  • Wan 2.2 T2V-A14B
  • Wan 2.2 I2V-A14B
  • StoryMem Wan2.2 M2V-A14B
  • +1
771 GitHub stars2 months agoCustom license

Trains and runs image, video and audio diffusion models on your own GPU

Model tooling
  • MiniMax H3
  • LTX-2.5
  • Wan-Animate-2
  • +7
13k GitHub starsyesterdayApache-2.0

Wan-Video

Wan2.2

Open video models for text, image, speech and character animation

Model tooling
  • Wan 2.2-T2V-A14B
  • Wan 2.2-I2V-A14B
  • Wan 2.2-TI2V-5B
  • +5
18k GitHub stars10 days agoApache-2.0

Featured on OpenMicroDrama

Do you maintain ShotStream? Add this badge to your README so English speakers can find our write-up.

Featured on OpenMicroDrama
markdown
[![Featured on OpenMicroDrama](https://openmicrodrama.com/badges/featured.svg)](https://openmicrodrama.com/projects/klingairesearch-shotstream)

The badge links to this page. Want your description changed? Email support@openmicrodrama.com.

Reviewed Oct 1, 2026. Reference implementation of an ECCV 2026 paper from MMLab CUHK and the Kling team; built on Wan2.1-T2V-1.3B. We write these descriptions ourselves. The repo's own docs are the final word.