Skip to content

Storyboarding a Micro Drama with AI: Script to Shot List to Video

Turn a micro drama script into a shot list, keyframes and storyboard panels, then image-to-video prompts. Plan blocking before you spend video credits.

Updated Oct 1, 2026

Plan each shot before you pay for video. If you skip that, the model decides where people stand, which way they face and how close the camera is. It decides differently every time, and you pay for every try.

A storyboard (a set of simple pictures that shows each shot) puts those choices back in your hands. Here's the order most AI drama creators use: script, shot list, keyframes, storyboard panels if you need them, then image-to-video.

Text-only prompts vs storyboard panels

You don't need to storyboard every shot. CottonByte AI explains the trade-off well in a detailed Seedance 2.5 tutorial. Use this table to pick an approach per scene.

ApproachBetter forWeak spot
Text-only shot descriptions (with character and location references)Dialogue scenes, emotional close-ups, demanding lighting. Usually better image qualityText describes a position without fixing it, so multi-character blocking takes more retries
Storyboard panels as a referenceSeveral people in frame, complex blocking, fights, batch productionThe board's look can leak into the final texture and style

Here's a simple rule for micro drama. Use text only for two-person close-up dialogue, which is most of an episode. Use boards for the gala fight, the chase, or any shot with three or more people.

Step 1: Turn the script into a shot list

Write one row per shot, with one action per row. Here's the paywall beat from our CEO romance template, broken into shots. (The paywall is the point where viewers have to pay to keep watching.)

#SecShotActionDialogueReferences
13InsertWall screen freezes on security footage of Elena photographing a filenoneboardroom plate
24Over-the-shoulderOver an investor's shoulder on the video call, Adrian at the head of the table"Is that your new assistant?" (V.O.)Adrian head + look A
33Close-upElena's hands go still on her notebooknoneElena look A
45MediumAdrian sets a velvet box on the glass table and slides it to her"She was photographing our engagement contract."Adrian, Elena, boardroom
54Close-upAdrian, low, eyes on the camera"Smile, fiancée. You're on camera."Adrian head
63Close-upElena doesn't smile. Cut to blacknoneElena head

The takeaway: six short shots, one action each, and every shot names the reference images it needs.

A chatbot can write the first draft of this table. Give it the scene, your character list and some hard rules:

Turn this scene into a vertical 9:16 shot list for AI video.
Rules: one action per shot; 3-8 seconds per shot; shot size from
[close-up, medium, wide, over-the-shoulder, tracking, establishing, insert];
max 2 speaking characters per shot; one line of dialogue per shot;
list which reference images each shot needs. Output a table.
Characters: ADRIAN VALE (...), ELENA BROOKS (...).
Scene: [paste]

Want something you can install instead? Our skills directory lists agent skills for Claude Code and Codex that turn scripts into storyboards. One example is the MIT-licensed zenstory drama-skills collection.

Step 2: Make keyframes with an image model

A keyframe is a still picture of the first frame of a shot. Make it at 9:16 before any video. Build it from your character pack and a picture of the location.

The creators we watched used different image models. Ryan Brooks and Isa used GPT Image 2, and Supercreator used Nano Banana inside Google Flow. They all work the same way: make a character sheet first, then point every frame at it.

Ryan Brooks makes one sheet per character. It has full-body front and back views plus front and side portraits. He adds that sheet to every storyboard prompt.

Keyframes help you twice. You catch mistakes, like the wrong outfit or the wrong side of the table, for the price of one image. And image-to-video (you give the model a still and it animates it) keeps the layout of the shot much better than text alone.

Step 3: Storyboard panels for blocking-heavy scenes

Blocking means where each person stands and moves. When people and the camera both move, a board with several panels shows the model where everyone goes. ByteDance's Seedance 2.5 prompt guide gives three rules:

  • Keep a multi-panel board to 15 panels or fewer.
  • Draw the panels as stick figures or simple line art.
  • Don't put text on the board.

CottonByte found the same thing in his tests. Seedance 2.5 copies references more closely than 2.0. So a detailed board can drag its sketch lines, hair and comic-panel layout into your video.

A simple board works better. Show the right number of people, where they stand, the key props and the camera direction. His point: a picture where the characters relate correctly beats a wrong one in 4K.

His other tips from the same video:

  • In a 30-second Seedance 2.5 clip, keep the main action inside the first 25 to 28 seconds. The end is where things start to drift (slowly change from what you asked for).
  • Need two shots to join up? Take the last frame of the first clip and make it black and white. Use it as the first-frame layout guide for the next clip.
  • Try a position map with colored mannequin figures and numbered spots. Then your prompt can say "the woman stands at yellow position 1." Tell the model not to draw the numbers.

Seedance 2.5 can also follow a simple 3D clay-model video for movement. ByteDance's guide says rough block shapes work better than detailed models.

Step 4: Write the image-to-video prompt per shot

Write each prompt in the same order so you don't forget anything. ByteDance's basic order for Seedance 2.5 is subject, action, scene, style, camera, then sound. That order works for most models.

[Image1 is Elena; Image2 is the boardroom. Use Image1 for face and outfit only.]
Subject: Elena Brooks, 24, chestnut low ponytail, cream silk blouse.
Action: her hands stop moving on her notebook; she looks up slowly.
Scene: glass boardroom, rainy city behind, investors on a wall screen.
Style: photoreal, soft overcast light, shallow depth of field.
Camera: close-up, slow push-in, 9:16.
Sound: rain on glass, a speakerphone voice fading out. No subtitles.

Each model handles multiple shots a little differently:

  • Kling 3.0 can split one clip into several shots with its multi-prompt option.
  • Seedance 2.0: ByteDance's guide says to use "Shot 1 / Shot 2" labels. It warns that exact timing like "0–3 seconds" isn't reliable.
  • Seedance 2.5 takes either timestamps or shot labels.

Our prompt library has ready-made prompts for each model: Seedance, Kling, Veo, Hailuo and Wan.

Step 5: Watch a rough cut before you pay for finals

Make cheap drafts first. Seedance 2.5's API has a draft mode that renders a 480p preview. It gives you an ID you can finish at 1080p within seven days.

With other models, render at the lowest resolution.

Cut the drafts together in order and watch the episode on your phone. Then remake only the shots that are broken. In Isa's OpenArt tutorial, three of five shots came back wrong, so she remade just those three with the same cast and settings.

Tools

Open-source tools are growing fast. Many are Chinese-first projects, and their English docs lag behind. These four are worth knowing:

  • Seedance2-Storyboard-Generator turns a novel or story into a multi-episode script and storyboard prompts for Seedance 2.0.
  • zenstory drama-skills (MIT) is a skill pack for Claude Code and Codex. It covers script, character assets, storyboards and prompts.
  • LocalMiniDrama (MIT) runs story, storyboard and video steps on your own computer.
  • Toonflow (MIT) is a canvas-style platform with agents and node workflows, and it includes storyboarding.

We track these and more, with English summaries, in storyboard projects and full production platforms. Paid canvas tools like Google Flow and OpenArt are in the tools directory.

Before you generate video

  • Every shot has one action and one shot size.
  • Every shot lists its reference images.
  • Dialogue is one line per shot, in one language.
  • Scenes with lots of movement have a simple board or position map.
  • The rough cut makes sense when you watch it on a phone.

Next, learn how the vertical frame works in vertical framing and shot language. Or follow the whole process in how to make an AI micro drama.

Sources

  1. docs.byteplus.com/en/docs/ModelArk/2607689
  2. docs.byteplus.com/en/docs/ModelArk/2607688
  3. docs.byteplus.com/en/docs/ModelArk/2222480
  4. fal.ai/models/bytedance/seedance-2.5/reference-to-video
  5. fal.ai/models/fal-ai/kling-video/v3/pro/image-to-video
  6. www.youtube.com/watch?v=r9Z4GC7nb_8
  7. www.youtube.com/watch?v=96q3KHK9KEE
  8. www.youtube.com/watch?v=XeWUtjzDIHw
  9. www.youtube.com/watch?v=vIksC7naeHo
  10. www.youtube.com/watch?v=dOmKYJoRboE
  11. github.com/zenstory-ai/drama-skills
  12. github.com/liangdabiao/Seedance2-Storyboard-Generator
  13. github.com/xuanyustudio/LocalMiniDrama
  14. github.com/HBAI-Ltd/Toonflow-app