Plan your sound in four layers, and keep the same music kit for the whole series. Sound is half of a micro drama. Mute a ReelShort-style episode and it plays like a slideshow.
The swell under the confession, the "dun" when the CEO walks in, the phone buzz at the end: that's what makes it feel like a show. AI video gives you the pictures. The sound pass does the rest.
The four layers
Build every episode from these four layers. They're listed from most to least important.
- Dialogue. Keep it clean, steady and always on top. See lip-sync and dubbing.
- Ambience and room tone. Ambience is the background sound of a place, like rain on glass or crickets outside. Room tone is the quiet hum of a room, and it glues shots from different clips together.
- SFX and Foley. SFX means sound effects. Foley is everyday sounds like footsteps, a glass set down too hard or a door slam, and they make the AI picture feel solid.
- Music. Use a score bed (quiet music under a scene). Add short stingers (one-second music hits) on reveals and cliffhangers.
Native audio: keep some, mute some
Native audio means the model makes sound along with the picture. Veo 3.1, Seedance 2.x, Kling 3.0, MiniMax H3 and recent Wan versions can do this. Use it, but pick what to keep.
Usually keep: ambience and the obvious sounds in the shot. They're already in sync.
Usually replace: the music. Each clip makes up its own score, so a scene cut from six clips gets six different songs. The cuts feel jumpy, and you can't build a theme.
Ask for sound the way each model expects:
- Veo. Its docs use labels like "SFX:" and "Ambient noise:". Put dialogue in quotation marks.
- MiniMax H3. Its prompt guide splits the overall soundscape from non-diegetic music (music the characters can't hear). That makes it easy to ask for ambience with no score.
- Seedance. Its prompt guide says negative wording works for audio, so "no BGM" (no background music) is fine. Its FAQ says to avoid water, wave and echo instructions, which cause bubbly sound, and to fade out clip audio at joins to remove clicks.
- Silent models, like Hailuo 2.3. These need a full sound pass. One creator calls it building a "sound skeleton" before the edit: rough ambience and hits first, picture timing second.
Model details are in our guides for Veo, Seedance, Kling, Hailuo and Wan.
Build a stinger kit
Make 5 to 8 music cues per series. Reuse them in all 60 to 100 episodes. Viewers learn them fast.
- The hit: one low impact for the slap, the reveal or the ring box opening.
- The riser: 2 to 4 seconds of building tension into the episode's last line.
- The cut to black: a hard stop or a reverse cymbal on the cliffhanger frame.
- The romance swell: strings or soft pads for the almost-kiss.
- The comedy sting: a pluck or a record scratch for the reaction shot.
- The theme: 10 to 15 seconds you can trim to open or close episodes.
A kit does two jobs. You don't have to make new music every episode, and people know your show three seconds in. Our script templates mark a cliffhanger every 60 to 90 seconds, so put your stingers there.
Do the sound pass in this order
This is the order ElevenLabs showed for scoring an AI short. It works for drama too.
- Lock the picture first. Re-timing music after every edit wastes hours.
- Lay the music bed under the whole scene so you can feel the pace.
- Add SFX. Describe what's on screen and a bit beyond it: steam, an engine starting, a flickering bulb, low machine hum.
- Balance. Background sounds go low. The one sound that matters goes up.
- Duck the music under dialogue. Ducking means turning the music down while someone talks. If you can't hear a word on a phone speaker, the music is too loud.
- Export a separate dialogue stem. A stem is a file with just the voices. You'll want it for dubbing later.
Music tools and what they let you do
Your rights depend on the plan you were on when you made the track. The plan you have now doesn't matter. Check the current terms before you publish.
| Tool | Plan and price (as of Oct 1, 2026) | Commercial use |
|---|---|---|
| Suno Free | $0, 50 credits a day | No commercial rights, no monthly downloads |
| Suno Pro | $8/month billed yearly, 2,500 credits, 20 song downloads a month | Yes, for songs made while subscribed |
| Suno Premier | $24/month billed yearly, 10,000 credits, 60 downloads a month | Yes, for songs made while subscribed |
| ElevenLabs Music | Credits from the shared pool, about 900 per minute of music | Music commercial use listed from the $6 Starter plan |
| ElevenLabs Sound Effects | About 200 credits per generation | Commercial license from Starter |
| Udio | Check current status | Downloads are restricted inside a "walled garden" |
The takeaway: for music you can sell, you need a paid Suno or ElevenLabs plan when you make the track.
Suno's help center says paying later doesn't give you commercial rights to songs you made on the free plan. Udio has kept songs inside a "walled garden" since its October 2025 deal with Universal Music Group, while it builds a licensed platform. So for now, you can't really use it for episode music you need to export.
Open-source and free sources
| Source | License | Use it for |
|---|---|---|
| ACE-Step 1.5 | MIT | Local music generation, including on consumer GPUs |
| YuE2 | Weights CC-BY-NC-4.0 | Songs with vocals, but non-commercial |
| MMAudio | Code MIT, checkpoints CC-BY-NC 4.0 | Video-to-audio Foley, non-commercial |
| HunyuanVideo-Foley | Tencent Hunyuan Community License | Video-to-audio Foley; the license doesn't apply in the EU, UK or South Korea |
| YouTube Audio Library | YouTube Audio Library license or Creative Commons per track | Copyright-safe music and SFX for YouTube videos |
| Freesound | Creative Commons, set per sound | Field recordings and SFX; check each file's license |
The takeaway: read the license column first. Anything marked NC (non-commercial) is for tests, not for episodes you sell.
YouTube says Audio Library tracks won't get Content ID claims (YouTube's automatic copyright matches). You can also monetize them in the Partner Program. Creative Commons tracks still need the artist credited in your description, and if you post beyond YouTube, read each track's license first.
Keep a sound log
For every cue in every episode, write down the source, the plan or license, the date and a link. It takes a minute per episode. It saves you when a platform sends a claim or a distributor asks for proof of your rights, as covered in publishing and monetizing.
Set the loudness
Make every episode equally loud. Platforms turn loud uploads down and quiet ones up, so uneven episodes just sound inconsistent. Loudness is measured in LUFS (an average loudness level; closer to zero is louder).
Spotify publishes its numbers: it normalizes to -14 LUFS integrated and leaves 1 dB of headroom for lossy encodes. Video platforms don't all publish targets. Use -14 as a sensible reference and keep every episode the same.
FFmpeg's loudnorm filter does EBU R128 normalization, a broadcast loudness standard. One pass on a finished episode looks like this:
ffmpeg -i ep01.mp4 -af loudnorm=I=-14:TP=-1:LRA=11 -c:v copy ep01_norm.mp4
Run it on every episode with the same settings. Then listen on a phone speaker and on earbuds. Dialogue should sit clearly above the music on both.
Quick checklist
- Native music muted on every clip, ambience kept where it matches
- One stinger kit, reused across the series
- Music ducked under every line
- Clicks removed at clip joins (short fades)
- Sound log updated
- Episode normalized with the same loudness settings as the last one
Next, read editing and captions to cut the episode together. Or see /tools for the commercial audio apps we track.
Sources
- suno.com/pricing
- help.suno.com/en/categories/550145-rights-ownership
- elevenlabs.io/pricing
- www.universalmusic.com/universal-music-group-and-udio-announce-udios-first-strategic-agreements-for-new-licensed-ai-music-creation-platform/
- www.digitalmusicnews.com/2026/07/17/udio-buydrm-deal-walled-garden/
- github.com/ace-step/ACE-Step-1.5
- huggingface.co/m-a-p/YuE2-3B
- github.com/hkchengrex/MMAudio
- github.com/Tencent-Hunyuan/HunyuanVideo-Foley
- support.google.com/youtube/answer/3376882?hl=en
- freesound.org/help/faq/
- support.spotify.com/us/artists/article/loudness-normalization/
- ffmpeg.org/ffmpeg-filters.html#loudnorm
- ai.google.dev/gemini-api/docs/veo
- docs.byteplus.com/en/docs/ModelArk/2607689
- huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
- www.youtube.com/watch?v=-k3a94vJJ1I
- www.youtube.com/watch?v=TVtr1WsBaEk