Skip to content

Self-Hosting an Open-Source AI Micro-Drama Pipeline

Build a self-hosted AI micro-drama pipeline: open drama platforms, open-weight video models with real VRAM numbers, TTS, lip-sync, and the licenses that matter.

Updated Oct 1, 2026

Run the app and your files yourself, and keep paying an API (a pay-per-clip cloud service) only for the final video shots. That's the setup most small teams end up with. Self-hosting means running the software on your own computer or a rented one.

Why bother? Per-second fees add up when a series is 80 episodes and every shot takes three tries. Self-hosting swaps that bill for a different one: a GPU (graphics card), setup time and a quality gap you'll have to manage.

Two honest warnings before you buy hardware. The best closed models (Seedance, Kling, Veo) can't be downloaded, so a fully local setup usually trails them on faces and motion. And "open source" covers many licenses, and some don't allow commercial use at all.

The pipeline at a glance

Every self-hosted setup follows these same steps. You can run some of them locally and send others to the cloud.

Script & episode outline (LLM)
  → Scene breakdown and shot list
  → Character sheets + keyframes (image model)
  → Video clips per shot (Wan / LTX / Hunyuan / H3)
  → Voices (TTS or actors) → lip-sync on dialogue shots
  → Edit, captions, music, loudness → export 9:16

You don't have to self-host all of it. Most creators who "self-host" run the app and files locally. They still call cloud APIs for the hardest step, the video.

Option A: run an open drama platform

A drama platform is one app that runs the whole flow: script, characters, storyboard, clips and the stitched episode. You run the app, then plug in API keys or local models. We track these under platform projects.

ProjectLicenseWhat stands out
ToonflowMITInfinite canvas for script, assets and clips; desktop, Docker or server; connects third-party APIs or local ComfyUI and LLMs
LocalMiniDramaMITLocal-first desktop app (SQLite and local files); story to episode with 9:16 support; local Ollama for text
JellyfishApache-2.0Production workspace built around consistency of characters, scenes, props and costumes; Docker Compose
LumenXMITStudio pipeline plus a playground; CosyVoice and Qwen3-TTS voices; FFmpeg export
huobao-dramaCC BY-NC-SA 4.0Full-stack TypeScript with agent roles (script rewriter, extractor); Docker image. Non-commercial license
ArcReelAGPL-3.0Novel or script to episodes with reusable assets, per-shot redo, cost tracking; Docker Compose; exports Jianying drafts (CapCut compatibility not yet verified, per the README)
OpenMontageAGPL-3.0General agentic video production; lists local Wan and Hunyuan through ComfyUI
dramaclawElastic License 2.0Local MCP server for Claude Code and Codex; gateway can add a local ComfyUI workflow
waoowaooElastic License 2.0Canvas workspace with an assistant; self-hosted preview via Docker; models through OpenRouter; no music or voiceover yet

The takeaway: Toonflow, LocalMiniDrama, Jellyfish and LumenX have the easiest licenses (MIT or Apache-2.0) if you plan to sell your series.

Toonflow's README records one real case. It took about two hours to make a roughly 2-minute piece with Seedance 2.0, GPT Image 2 and Claude Opus 4.6. Model calls cost about ¥130, and ¥120 of that was video.

That's one data point, not a promise. But it shows where the money goes: video, by far.

Read the license before you build a business on it

  • MIT and Apache-2.0: you can use, change and sell it. Just give credit.
  • AGPL-3.0: if you change it and let others use it over a network, you must share your source code.
  • CC BY-NC-SA 4.0: no commercial use. It's fine for learning, not for a paid series.
  • Elastic License 2.0: you can read the code, but it isn't open source by the OSI definition (the standard list of approved open licenses). Among other limits, you can't offer it to others as a hosted service.

This is a summary, not legal advice. The license file in each repo is what counts.

Option B: run the models yourself

Video

Open weights means you can download the model and run it yourself. VRAM is your graphics card's memory, and it decides which models you can run.

ModelLicenseHardware notes from the official docs
Wan 2.2 TI2V-5BApache-2.0720p at 24 fps; README says it runs on a 24 GB card like the RTX 4090; ComfyUI's docs say the 5B "should fit well on 8GB vram" with native offloading
Wan 2.2 A14B (T2V, I2V)Apache-2.0README single-GPU commands need 80 GB VRAM; offload and dtype flags cut memory
Wan 2.2 S2V-14B and Animate-14BApache-2.0Speech-to-video, and character animation or replacement
Wan-Animate-2Apache-2.0Newer character animation weights (August 2026)
LTX-2.xLTX-2.x Community LicenseAudio and video in one pass; ComfyUI lists LTX-2 as a 19B model. Organizations with $10M+ annual revenue need a paid license
HunyuanVideo 1.5Tencent Hunyuan Community License8.3B parameters; 14 GB minimum with offloading per the README. License doesn't apply in the EU, UK or South Korea
MiniMax H3MiniMax H3 Community LicenseNative stereo audio, 4 to 15 second clips. The license's territory excludes the EU, UK, South Korea and the US, which need a formal license application

The takeaway: Wan 2.2 has the friendliest license and the widest range of hardware. Check the territory rules for HunyuanVideo and MiniMax H3 before you start.

Offloading means parking parts of the model in normal memory when the graphics card runs out. It's slower, but it lets smaller cards run bigger models.

Wan versions after 2.2 (2.5, 2.6, 2.7, 3.0) are API-only. See our Wan guide for what each one adds.

Here's what creators report in practice. A pixaroma tutorial says some Wan 2.2 setups run on 8 GB with GGUF quantized models (smaller, compressed versions of the model). Heavier workflows like InfiniteTalk need much more; he tested on 24 GB and found some workflows work on 16 GB.

Another creator ran MiniMax H3 with a Turbo LoRA in ComfyUI on a 16 GB GPU. A LoRA is a small add-on file that changes how a model behaves, here to make it faster. It took about five minutes per short clip, and Wan2GP exists just for low-VRAM machines.

Everything else

StepOpen optionLicense
Keyframes and character sheetsQwen-ImageApache-2.0
Keyframes (alternative)FLUX.1 devNon-commercial license; check before commercial use
VoicesFun-CosyVoice 3, ChatterboxApache-2.0, MIT
Lip-syncLatentSync, InfiniteTalkApache-2.0
MusicACE-Step 1.5MIT
Character LoRA trainingmusubi-tuner (HunyuanVideo, Wan 2.1/2.2), ai-toolkit (Wan 2.1/2.2, LTX-2.x, MiniMax H3)Check repo; ai-toolkit is MIT
Edit and assemblyFFmpegLGPL/GPL

The takeaway: every other step has an open option too. Keyframes are images you make first and then animate, and TTS (text-to-speech) tools like CosyVoice turn written lines into voices.

For more on voices and lip-sync (matching mouth movements to words), see lip-sync and dubbing. Browse repos by step under workflow, model tooling, editing and video frameworks.

Hardware tiers, roughly

These come from the numbers above, not from our own tests:

  • 8 GB: Wan 2.2 5B in ComfyUI with offloading, and LatentSync 1.5. Expect slow, short, low-res drafts.
  • 16 GB: some Wan workflows, and H3 with a speed LoRA, per one creator's report.
  • 24 GB (RTX 4090 class): Wan 2.2 5B at 720p per the README, HunyuanVideo 1.5 per ComfyUI's docs, and most lip-sync tools.
  • 80 GB (data-center GPUs): Wan 2.2 A14B at full size on a single card.

Don't own the card? Rent one by the hour from a cloud GPU host. Shut it off between sessions.

A hybrid that works

Small teams most often use this setup:

  1. Run a platform locally, like Toonflow or LocalMiniDrama. Your scripts, character sheets and every take stay in your own folders.
  2. Draft shots on a local open model. Settle framing and timing for free.
  3. Send only each shot's final version to a paid API model.
  4. Do voices, lip-sync, music and the edit locally.

You only pay for the frames viewers actually see. To price it out, use the cost breakdown. To pick the paid model, see choosing a video model.

Agents and skills

Several platforms now offer MCP servers or skill files. MCP is a standard way for an AI agent like Claude Code or Codex to use outside tools. With it, the agent can write the episode, pull out characters, build the shot list and queue the clips.

We collect those under /skills. If you'd rather wire things step by step, start with ComfyUI for micro drama.

Sources

  1. github.com/HBAI-Ltd/Toonflow-app
  2. github.com/chatfire-AI/huobao-drama
  3. github.com/xuanyustudio/LocalMiniDrama
  4. github.com/Forget-C/Jellyfish
  5. github.com/alibaba/lumenx
  6. github.com/ArcReel/ArcReel
  7. github.com/waooAI/waoowaoo
  8. github.com/dramaclaw/dramaclaw
  9. github.com/calesthio/OpenMontage
  10. github.com/Wan-Video/Wan2.2
  11. github.com/Wan-Video/Wan-Animate-2
  12. docs.comfy.org/tutorials/video/wan/wan2_2
  13. github.com/Lightricks/LTX-2
  14. docs.comfy.org/tutorials/video/ltx/ltx-2
  15. github.com/Tencent-Hunyuan/HunyuanVideo-1.5
  16. docs.comfy.org/tutorials/video/hunyuan/hunyuan-video-1-5
  17. huggingface.co/MiniMaxAI/MiniMax-H3
  18. huggingface.co/Qwen/Qwen-Image
  19. huggingface.co/black-forest-labs/FLUX.1-dev
  20. github.com/QwenAudio/CosyVoice
  21. github.com/resemble-ai/chatterbox
  22. github.com/bytedance/LatentSync
  23. github.com/MeiGen-AI/InfiniteTalk
  24. github.com/ace-step/ACE-Step-1.5
  25. github.com/kohya-ss/musubi-tuner
  26. github.com/ostris/ai-toolkit
  27. github.com/deepbeepmeep/Wan2GP
  28. ffmpeg.org/legal.html
  29. www.youtube.com/watch?v=Z8JlJdXdVg4
  30. www.youtube.com/watch?v=Z8giDB2wzFc