Run the app and your files yourself, and keep paying an API (a pay-per-clip cloud service) only for the final video shots. That's the setup most small teams end up with. Self-hosting means running the software on your own computer or a rented one.
Why bother? Per-second fees add up when a series is 80 episodes and every shot takes three tries. Self-hosting swaps that bill for a different one: a GPU (graphics card), setup time and a quality gap you'll have to manage.
Two honest warnings before you buy hardware. The best closed models (Seedance, Kling, Veo) can't be downloaded, so a fully local setup usually trails them on faces and motion. And "open source" covers many licenses, and some don't allow commercial use at all.
The pipeline at a glance
Every self-hosted setup follows these same steps. You can run some of them locally and send others to the cloud.
Script & episode outline (LLM)
→ Scene breakdown and shot list
→ Character sheets + keyframes (image model)
→ Video clips per shot (Wan / LTX / Hunyuan / H3)
→ Voices (TTS or actors) → lip-sync on dialogue shots
→ Edit, captions, music, loudness → export 9:16
You don't have to self-host all of it. Most creators who "self-host" run the app and files locally. They still call cloud APIs for the hardest step, the video.
Option A: run an open drama platform
A drama platform is one app that runs the whole flow: script, characters, storyboard, clips and the stitched episode. You run the app, then plug in API keys or local models. We track these under platform projects.
| Project | License | What stands out |
|---|---|---|
| Toonflow | MIT | Infinite canvas for script, assets and clips; desktop, Docker or server; connects third-party APIs or local ComfyUI and LLMs |
| LocalMiniDrama | MIT | Local-first desktop app (SQLite and local files); story to episode with 9:16 support; local Ollama for text |
| Jellyfish | Apache-2.0 | Production workspace built around consistency of characters, scenes, props and costumes; Docker Compose |
| LumenX | MIT | Studio pipeline plus a playground; CosyVoice and Qwen3-TTS voices; FFmpeg export |
| huobao-drama | CC BY-NC-SA 4.0 | Full-stack TypeScript with agent roles (script rewriter, extractor); Docker image. Non-commercial license |
| ArcReel | AGPL-3.0 | Novel or script to episodes with reusable assets, per-shot redo, cost tracking; Docker Compose; exports Jianying drafts (CapCut compatibility not yet verified, per the README) |
| OpenMontage | AGPL-3.0 | General agentic video production; lists local Wan and Hunyuan through ComfyUI |
| dramaclaw | Elastic License 2.0 | Local MCP server for Claude Code and Codex; gateway can add a local ComfyUI workflow |
| waoowaoo | Elastic License 2.0 | Canvas workspace with an assistant; self-hosted preview via Docker; models through OpenRouter; no music or voiceover yet |
The takeaway: Toonflow, LocalMiniDrama, Jellyfish and LumenX have the easiest licenses (MIT or Apache-2.0) if you plan to sell your series.
Toonflow's README records one real case. It took about two hours to make a roughly 2-minute piece with Seedance 2.0, GPT Image 2 and Claude Opus 4.6. Model calls cost about ¥130, and ¥120 of that was video.
That's one data point, not a promise. But it shows where the money goes: video, by far.
Read the license before you build a business on it
- MIT and Apache-2.0: you can use, change and sell it. Just give credit.
- AGPL-3.0: if you change it and let others use it over a network, you must share your source code.
- CC BY-NC-SA 4.0: no commercial use. It's fine for learning, not for a paid series.
- Elastic License 2.0: you can read the code, but it isn't open source by the OSI definition (the standard list of approved open licenses). Among other limits, you can't offer it to others as a hosted service.
This is a summary, not legal advice. The license file in each repo is what counts.
Option B: run the models yourself
Video
Open weights means you can download the model and run it yourself. VRAM is your graphics card's memory, and it decides which models you can run.
| Model | License | Hardware notes from the official docs |
|---|---|---|
| Wan 2.2 TI2V-5B | Apache-2.0 | 720p at 24 fps; README says it runs on a 24 GB card like the RTX 4090; ComfyUI's docs say the 5B "should fit well on 8GB vram" with native offloading |
| Wan 2.2 A14B (T2V, I2V) | Apache-2.0 | README single-GPU commands need 80 GB VRAM; offload and dtype flags cut memory |
| Wan 2.2 S2V-14B and Animate-14B | Apache-2.0 | Speech-to-video, and character animation or replacement |
| Wan-Animate-2 | Apache-2.0 | Newer character animation weights (August 2026) |
| LTX-2.x | LTX-2.x Community License | Audio and video in one pass; ComfyUI lists LTX-2 as a 19B model. Organizations with $10M+ annual revenue need a paid license |
| HunyuanVideo 1.5 | Tencent Hunyuan Community License | 8.3B parameters; 14 GB minimum with offloading per the README. License doesn't apply in the EU, UK or South Korea |
| MiniMax H3 | MiniMax H3 Community License | Native stereo audio, 4 to 15 second clips. The license's territory excludes the EU, UK, South Korea and the US, which need a formal license application |
The takeaway: Wan 2.2 has the friendliest license and the widest range of hardware. Check the territory rules for HunyuanVideo and MiniMax H3 before you start.
Offloading means parking parts of the model in normal memory when the graphics card runs out. It's slower, but it lets smaller cards run bigger models.
Wan versions after 2.2 (2.5, 2.6, 2.7, 3.0) are API-only. See our Wan guide for what each one adds.
Here's what creators report in practice. A pixaroma tutorial says some Wan 2.2 setups run on 8 GB with GGUF quantized models (smaller, compressed versions of the model). Heavier workflows like InfiniteTalk need much more; he tested on 24 GB and found some workflows work on 16 GB.
Another creator ran MiniMax H3 with a Turbo LoRA in ComfyUI on a 16 GB GPU. A LoRA is a small add-on file that changes how a model behaves, here to make it faster. It took about five minutes per short clip, and Wan2GP exists just for low-VRAM machines.
Everything else
| Step | Open option | License |
|---|---|---|
| Keyframes and character sheets | Qwen-Image | Apache-2.0 |
| Keyframes (alternative) | FLUX.1 dev | Non-commercial license; check before commercial use |
| Voices | Fun-CosyVoice 3, Chatterbox | Apache-2.0, MIT |
| Lip-sync | LatentSync, InfiniteTalk | Apache-2.0 |
| Music | ACE-Step 1.5 | MIT |
| Character LoRA training | musubi-tuner (HunyuanVideo, Wan 2.1/2.2), ai-toolkit (Wan 2.1/2.2, LTX-2.x, MiniMax H3) | Check repo; ai-toolkit is MIT |
| Edit and assembly | FFmpeg | LGPL/GPL |
The takeaway: every other step has an open option too. Keyframes are images you make first and then animate, and TTS (text-to-speech) tools like CosyVoice turn written lines into voices.
For more on voices and lip-sync (matching mouth movements to words), see lip-sync and dubbing. Browse repos by step under workflow, model tooling, editing and video frameworks.
Hardware tiers, roughly
These come from the numbers above, not from our own tests:
- 8 GB: Wan 2.2 5B in ComfyUI with offloading, and LatentSync 1.5. Expect slow, short, low-res drafts.
- 16 GB: some Wan workflows, and H3 with a speed LoRA, per one creator's report.
- 24 GB (RTX 4090 class): Wan 2.2 5B at 720p per the README, HunyuanVideo 1.5 per ComfyUI's docs, and most lip-sync tools.
- 80 GB (data-center GPUs): Wan 2.2 A14B at full size on a single card.
Don't own the card? Rent one by the hour from a cloud GPU host. Shut it off between sessions.
A hybrid that works
Small teams most often use this setup:
- Run a platform locally, like Toonflow or LocalMiniDrama. Your scripts, character sheets and every take stay in your own folders.
- Draft shots on a local open model. Settle framing and timing for free.
- Send only each shot's final version to a paid API model.
- Do voices, lip-sync, music and the edit locally.
You only pay for the frames viewers actually see. To price it out, use the cost breakdown. To pick the paid model, see choosing a video model.
Agents and skills
Several platforms now offer MCP servers or skill files. MCP is a standard way for an AI agent like Claude Code or Codex to use outside tools. With it, the agent can write the episode, pull out characters, build the shot list and queue the clips.
We collect those under /skills. If you'd rather wire things step by step, start with ComfyUI for micro drama.
Sources
- github.com/HBAI-Ltd/Toonflow-app
- github.com/chatfire-AI/huobao-drama
- github.com/xuanyustudio/LocalMiniDrama
- github.com/Forget-C/Jellyfish
- github.com/alibaba/lumenx
- github.com/ArcReel/ArcReel
- github.com/waooAI/waoowaoo
- github.com/dramaclaw/dramaclaw
- github.com/calesthio/OpenMontage
- github.com/Wan-Video/Wan2.2
- github.com/Wan-Video/Wan-Animate-2
- docs.comfy.org/tutorials/video/wan/wan2_2
- github.com/Lightricks/LTX-2
- docs.comfy.org/tutorials/video/ltx/ltx-2
- github.com/Tencent-Hunyuan/HunyuanVideo-1.5
- docs.comfy.org/tutorials/video/hunyuan/hunyuan-video-1-5
- huggingface.co/MiniMaxAI/MiniMax-H3
- huggingface.co/Qwen/Qwen-Image
- huggingface.co/black-forest-labs/FLUX.1-dev
- github.com/QwenAudio/CosyVoice
- github.com/resemble-ai/chatterbox
- github.com/bytedance/LatentSync
- github.com/MeiGen-AI/InfiniteTalk
- github.com/ace-step/ACE-Step-1.5
- github.com/kohya-ss/musubi-tuner
- github.com/ostris/ai-toolkit
- github.com/deepbeepmeep/Wan2GP
- ffmpeg.org/legal.html
- www.youtube.com/watch?v=Z8JlJdXdVg4
- www.youtube.com/watch?v=Z8giDB2wzFc