Spend real time on the edit. It's where your AI clips start to feel like a show, and it's the cheapest step in the whole process.
Raw clips come back with slow starts and dead air at the end. The color shifts a little from shot to shot, and there are no captions. A tight edit fixes most of that.
Pick an editor
Any timeline editor works. These are the ones AI drama creators used most in the tutorials we watched. Use the table to pick one that fits your budget.
| Editor | Cost | Why people use it for micro drama |
|---|---|---|
| CapCut | Free version plus paid subscription plans; prices vary by country and app store | Fast vertical templates, auto captions, big sound effects library. Isa does AI and CottonByte AI both finish in CapCut |
| DaVinci Resolve | Free; Resolve Studio is a one-time $295 | The free version edits and finishes up to Ultra HD at up to 60 fps. Strong color tools for matching AI clips |
| Adobe Premiere | Subscription, see Adobe's plans page | Speech to Text for auto transcripts and captions, safe-zone overlay presets from the community |
The takeaway: you can do the whole job for free in CapCut or DaVinci Resolve.
Prices come from each company's site as of October 2026, so check before you buy. You'll find more options in the tools directory.
Set the timeline up right
Make a 1080 by 1920 vertical timeline before you import anything. Then check frame rates (frames per second, or fps), because AI models don't all use the same one.
For example, Alibaba's docs list 30 fps for Wan 3.0. fal's Seedance pricing formula assumes 24 frames per second. Pick one rate for the whole series and let the editor convert the rest.
YouTube's upload advice is to export at the same frame rate the video was made in.
Keep a project template for each series. Put the timeline, caption style, title card, end card and music bed in it. Then episode 47 takes minutes to put together, not an afternoon.
Cutting for pace
Vertical drama moves faster than people expect. These habits keep episodes tight:
- Trim the start and end of every clip. Many clips ease in slowly and drift at the end. Cut to the first frame where something happens.
- Smooth the joins. For joined or extended Seedance clips, ByteDance suggests trimming 6 frames from the end of the first clip and 1 frame from the start of the next.
- Cut on action. Cut on a door slam, a turn, or a hand hitting the table. Movement hides small differences between AI shots.
- Hold the reaction a beat longer than the line. The face after the reveal is the moment people screenshot.
- Let the sound lead. Start the next shot's sound a few frames early (editors call this a J-cut). It makes AI shots feel connected, even when you made them minutes apart.
- End on the hook line, then cut to black. Cut on the question. Don't let a clip drift and soften the cliffhanger.
- Match color last. Do a quick color pass so skin tones match across the episode. The free version of Resolve is plenty for this.
If a shot doesn't work, remake that shot, not the whole scene. In her OpenArt walkthrough, Isa fixes episodes this way: she remakes three bad shots and drops them back in place.
Captions: assume the sound is off
Burn your captions into the video. In her micro drama tutorial, Isa says that for vertical feeds you really want captions, because so many people watch with the sound off.
Don't rely on the platform's auto-captions. They're hit and miss, and viewers can turn them off.
What works on a phone:
- One or two short lines. Break lines where you'd pause when speaking. Long lines make people read instead of watching faces.
- Big, bold, plain letters with a shadow or outline. Caption tutorials often suggest the fonts Montserrat or Poppins. A soft shadow keeps text readable over bright backgrounds.
- Put them in the middle band. Keep captions above the app's buttons at the bottom and away from the right edge. See vertical framing for the safe-zone numbers.
- Use speaker colors sparingly. One accent color for the lead's lines is usually enough.
- Keep animation light. Word-by-word pop-up captions suit comedy and confessions. For tense scenes, plain still lines read better.
- Caption the sounds that matter. A gunshot, a door, a phone buzzing. Viewers with the sound off miss the cue otherwise.
Don't let the video model write captions. Several models add unwanted subtitles to vertical shots, and they're often misspelled. Put "no subtitles" in your prompts and add text in the edit.
Subtitle tools
CapCut's auto captions and Premiere's Speech to Text are fine for one episode at a time. For a full season, or for translations, free Whisper tools are faster because you can run many files at once. Pick from this table based on what you need.
| Tool | License | What it adds |
|---|---|---|
| OpenAI Whisper | MIT | The original speech recognition model and CLI. Outputs SRT and VTT |
| faster-whisper | MIT | A faster reimplementation using CTranslate2 |
| WhisperX | BSD-2-Clause | Word-level timestamps and speaker diarization, useful for word-by-word captions |
| Subtitle Edit | MIT | A desktop editor for fixing timing and text by hand |
The takeaway: start with Whisper, and add WhisperX if you want word-by-word captions. (SRT and VTT are standard subtitle file types. Diarization means telling apart who is speaking.)
Here's a basic Whisper command that writes an SRT file:
whisper ep01-dialogue.wav --model turbo --language en --output_format srt \
--word_timestamps True --max_line_width 32 --max_line_count 2
The line-width options only work with word timestamps turned on. Export a dialogue-only audio track from your editor first, so music doesn't confuse it.
You wrote the script, so check the transcript against it and fix any names it misheard. Our editing projects page tracks more subtitle, dubbing and upscaling tools.
Export specs
Use the platform's own numbers, not guesses. This table shows what each help page says.
| Platform | What the help docs say |
|---|---|
| TikTok Studio (web upload) | MP4 or WebM, 720 by 1280 or higher, up to 30 minutes, under 10 GB |
| YouTube Shorts | Up to 3 minutes, square or vertical aspect ratio |
| YouTube (general upload settings) | MP4, H.264 high profile, AAC-LC or Opus audio at 48 kHz. About 8 Mbps recommended for 1080p at 24 to 30 fps |
For micro drama, export a 1080 by 1920 H.264 MP4 with AAC audio at 48 kHz. That works on every platform listed here.
Loudness is trickier. Neither TikTok's nor YouTube's upload help pages that we checked give a target.
Here's what to do instead. Use your editor's loudness meter and keep dialogue clearly above the music. Keep every episode at about the same level, so people watching back to back don't reach for the volume, and always check the final file on a phone speaker.
Music and sound design have their own guide: music and sound.
Covers and episode packaging
People need to recognize your series in a busy feed. These habits help:
- Pick a cover frame on purpose. TikTok Studio lets you choose a cover image when you upload. Use the hook frame: a close-up face mid-emotion gets more clicks than a wide shot.
- Keep the title style the same. Use the same font, position and episode number style on every cover, inside the safe zone.
- Number everything. Put "Ep 12" on the cover and in the caption. People watching back to back look for the next one.
- Add an end card to free episodes that points to where the rest lives. The publishing and monetizing guide covers where that should be.
Episode export checklist
- Clip starts and ends trimmed, joins smoothed.
- Color matched across shots.
- Captions burned in, inside the safe zone, and checked against the script for spelling.
- Dialogue at the same level as the last episode.
- Exported as a vertical H.264 MP4 at the series frame rate.
- Watched once on a real phone, with sound on and sound off.
If you need the voices first, see lip-sync and dubbing. If you're starting from zero, go to how to make an AI micro drama.
Sources
- www.blackmagicdesign.com/products/davinciresolve/studio
- www.capcut.com/help/how-much-does-capcut-pro-cost
- www.adobe.com/products/premiere/plans.html
- helpx.adobe.com/premiere/desktop/add-text-images/insert-captions/auto-transcribe-video-using-speech-to-text.html
- docs.byteplus.com/en/docs/ModelArk/2222480
- github.com/openai/whisper
- github.com/SYSTRAN/faster-whisper
- github.com/m-bain/whisperX
- github.com/SubtitleEdit/subtitleedit
- support.tiktok.com/en/using-tiktok/creating-videos/creator-tools-on-tiktok
- support.google.com/youtube/answer/1722171?hl=en
- support.google.com/youtube/answer/12779649?hl=en
- www.alibabacloud.com/help/en/model-studio/video-generation
- fal.ai/models/bytedance/seedance-2.5/text-to-video
- www.youtube.com/watch?v=vIksC7naeHo
- www.youtube.com/watch?v=dOmKYJoRboE
- www.youtube.com/watch?v=oUU1Lj7hcNU
- www.youtube.com/watch?v=r9Z4GC7nb_8