Skip to content

Veo 3.1 for AI Micro Dramas: Access, Pricing and 9:16 Best Practices

How to use Google Veo 3.1 for vertical micro dramas: which version to pick, Flow vs API, real per-second prices, dialogue, and keeping faces the same.

Google DeepMind · updated Oct 1, 2026

Use Veo when a scene has to talk. It makes the picture and the sound in one go. The actor's line, the room noise and the slammed door all come back together, with lip sync (mouth movements that match the words).

That matters for micro drama, because almost every beat is a close-up argument. There are two catches. One clip is 8 seconds at most, and Google now tells new projects to use a different model.

This page covers Veo 3.1 as of October 2026. You'll see which version to pick, where to run it and what a clip costs. You'll also see how to keep your cast looking the same across 80 episodes.

What Veo is good and bad at for micro drama

Good at

  • Dialogue from one prompt. You get lip-synced lines, sound effects and background sound together. Google's docs say to put speech in quotation marks and describe sounds clearly.
  • True vertical video. It makes 9:16 at 720p, 1080p or 4K. It's not a crop of a wide frame.
  • Film language. Google's prompt formula starts with the camera: shot size, camera move and lens.
  • Tools for continuity. You get up to three reference images (pictures you give the model so it copies a face or place). You also get first-and-last-frame control and Extend (adding more seconds to a clip).

Weaker at

  • Length. One generation is 4, 6 or 8 seconds. For anything longer, you extend or edit.
  • Two people talking in one clip. Creators report that Veo sometimes swaps or repeats lines between speakers.
  • Teen characters in image-based shots. Image-to-video (you give it a still and it animates it), frame interpolation and reference images only accept the "allow_adult" person setting. Google defines that as adults only.

Current versions

ModelGemini API IDStatusMax resolutionExtendReference images
Veo 3.1veo-3.1-generate-previewPreview; GA on Google Cloud since Nov 17, 20254KYesUp to 3
Veo 3.1 Fastveo-3.1-fast-generate-previewPreview; GA on Google Cloud4KYesUp to 3
Veo 3.1 Liteveo-3.1-lite-generate-previewPreview since Mar 31, 20261080pNoNo

The takeaway: pick Lite for cheap drafts, and Veo 3.1 or Fast when you need references or Extend. "GA" means generally available, so it's out of preview.

Veo 3.1 launched on October 15, 2025. Google shut down Veo 2, Veo 3 and Veo 3 Fast in the Gemini API on June 30, 2026. Older tutorials that use those model IDs won't run.

Test Gemini Omni Flash too. Google's video docs now say to use Gemini Omni Flash as the default video model. They keep Veo 3.1 for scene extension, last-frame control and existing pipelines.

Omni Flash 1.1 went GA on August 27, 2026. Flow's newest character and voice features run on it, but Flow lists Extend for Omni as "coming soon." Try both on one of your own scenes before you lock in a workflow.

Where to run it

  • Google Flow is Google's filmmaking app. It has Ingredients, Frames to Video, Extend and a scene builder, and you pay in Flow credits. Everyone gets 50 credits a day, and paid plans add more.
  • Flow model limits: Veo 3.1 Quality doesn't support ingredients or Extend in Flow. Fast supports ingredients, and Lite handles extends.
  • Gemini app: Pro includes a limited Veo 3.1 Lite trial, and Ultra includes Veo 3.1. It's fine for tests.
  • Gemini API (with a Google AI Studio key): you pay per second, and there's no free tier for Veo. An API means you send requests from code instead of using an app. It's best for scripted pipelines.
  • Google Cloud (Vertex AI, now part of Gemini Enterprise Agent Platform) has the GA model IDs. You get up to four videos per request, negative prompts (a list of things you don't want) and a cheaper video-only rate.
  • fal.ai and Replicate run the same models behind a simple API with an audio on/off switch. fal also has negative prompts and a safety-tolerance setting.

Pricing (as of October 2026)

Here are the per-second rates from each pricing page, worked out for one 8-second clip:

WhereVeo 3.1Veo 3.1 FastVeo 3.1 Lite
Gemini API, audio on$0.40/s → $3.20 (720p/1080p), $4.80 at 4K$0.10/s at 720p → $0.80; $0.12/s at 1080p → $0.96$0.05/s at 720p → $0.40; $0.08/s at 1080p → $0.64
Google Cloud, video only$0.20/s → $1.60$0.08–$0.10/s → $0.64–$0.80$0.03–$0.05/s → $0.24–$0.40
fal.ai, audio on$0.40/s → $3.20$0.15/s → $1.20$0.05–$0.08/s → $0.40–$0.64
Replicate, audio on$0.40/s → $3.20$0.15/s → $1.20$0.05–$0.08/s → $0.40–$0.64

The takeaway: at 720p with audio on the Gemini API, Fast costs a quarter of full Veo 3.1. Lite is cheaper still.

Flow credits: a Veo 3.1 Fast generation costs 20 credits (10 on Ultra), and a Quality clip costs 100. Google AI Pro ($19.99/month in the US) adds 1,000 credits a month. That's roughly 50 Fast clips or 10 Quality clips.

Ultra starts at $99.99 with 10,000 credits. Costs change, so check Flow's settings panel before a big batch. Set outputs per prompt to one so you don't pay for extras.

You aren't charged for blocked or failed generations on the API or in Flow. You are charged for rerolls, so budget for them. Our cost breakdown turns these rates into a budget per episode.

Best practices for 9:16 serialized drama

Lock the cast before you shoot

Make each lead once. Use a front-facing portrait and a full-body outfit shot on a plain background. Flow's help pages recommend plain or segmented backgrounds for references.

Also make an empty set photo for each place you keep coming back to. Give Veo up to three of these as reference images. Clips that use reference images must be 8 seconds long.

Then write one fixed description for each character. Paste it word for word into every prompt: "32-year-old CEO, sharp jaw, slicked-back black hair, charcoal three-piece suit, silver cufflinks." If you reword it between episodes, you'll quickly get a slightly different person.

Our character consistency guide has more.

Write your prompt like a shot list

Google's Veo 3.1 prompting guide uses a five-part formula: camera, subject, action, setting, then style and mood. Start with the shot size and camera move. Then add the character description and one clear action.

When a beat needs several cuts, use timestamps: [00:00-00:02] medium shot... then [00:02-00:04] close-up.... Four cuts inside one 8-second generation share the same cast and light. That usually looks better than stitching four separate clips.

Dialogue and lip sync

Put each line in quotes with a note on how it's said. For example: He says quietly, "You're late." Keep it to one speaker and one or two short lines per clip.

When two characters argue, make each side its own shot and cut between them. A TV editor would cover the scene that way anyway. It also avoids the swapped-lines problem.

Sometimes you'll get captions burned into the video. If so, list "subtitles, captions" as a negative prompt where you can (Google Cloud, fal). On the Gemini API, write "no on-screen text" in the prompt.

Our lip-sync and dubbing guide explains when to replace Veo's voices.

Clip length, Extend and stitching

Extend adds 7 seconds per step, up to 20 times, for about 148 seconds in total. It only works on 720p clips that Veo made. It continues from the final second, so it can't carry a voice unless that voice is speaking in the last second.

Creators report faces drifting (slowly changing) over long chains. So use Extend for walk-and-talks and slow moves. For each new angle, start a fresh generation locked to your references.

For a clean match between shots, use first-and-last-frame mode. Use the final frame of one clip as the first frame of the next. Download everything right away, because the Gemini API deletes generated videos after two days.

Pace for the scroll

Episodes win or lose viewers in the first three seconds. So open on a face or a reveal, not a slow skyline. Use 4- to 6-second clips for reaction cuts and keep 8-second clips for dialogue.

Keep faces in the upper middle of the frame so captions and app buttons don't cover them. See vertical framing and shot language and editing and captions.

Common mistakes

  • Asking for 1080p or 4K at 4 or 6 seconds. Those resolutions only work at 8 seconds, and Extend only works at 720p.
  • Expecting Lite to do everything. In the API, Lite has no reference images and no Extend.
  • Putting two speakers in one clip. Split them into a shot and a reverse shot.
  • Using reference images for teen characters. Image-based modes are adults only. Keep those shots text-to-video, or make the character older.
  • Writing negatives as "no walls." Google says to list the unwanted things instead: "wall, frame."
  • Treating a seed as a lock. A seed is a number that makes results repeatable. Google says it slightly helps but doesn't guarantee the same result.

Prompts and next steps

Start with our Veo prompt library. It has 40 original 9:16 prompts written the way Veo likes them, with quoted dialogue, labeled sound effects and timestamped multi-cut scenes.

Jump to CEO romance, revenge or thriller prompts. You can also pair them with the revenge script template.

Still picking a model? Compare Veo with Kling and Seedance in choosing a video model.

Sources

  1. deepmind.google/models/veo/
  2. ai.google.dev/gemini-api/docs/video
  3. ai.google.dev/gemini-api/docs/veo
  4. ai.google.dev/gemini-api/docs/pricing
  5. ai.google.dev/gemini-api/docs/changelog
  6. ai.google.dev/gemini-api/docs/omni
  7. cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1
  8. cloud.google.com/vertex-ai/generative-ai/pricing
  9. docs.cloud.google.com/gemini-enterprise-agent-platform/models/veo/3-1-generate
  10. docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/generate-videos-from-text
  11. docs.cloud.google.com/gemini-enterprise-agent-platform/models/video/video-gen-prompt-guide
  12. blog.google/technology/ai/veo-updates-flow/
  13. gemini.google/us/subscriptions/?hl=en
  14. support.google.com/labs/answer/16526234?hl=en
  15. support.google.com/labs/answer/16352836?hl=en
  16. support.google.com/labs/answer/16353334?hl=en
  17. fal.ai/models/fal-ai/veo3.1
  18. fal.ai/models/fal-ai/veo3.1/fast
  19. fal.ai/models/fal-ai/veo3.1/lite
  20. fal.ai/models/fal-ai/veo3.1/api
  21. replicate.com/google/veo-3.1
  22. replicate.com/google/veo-3.1-fast
  23. replicate.com/google/veo-3.1-lite
  24. www.youtube.com/watch?v=8TcXk7JJZ3g
  25. www.youtube.com/watch?v=VYH-zwR8iQU
  26. www.youtube.com/watch?v=c772SYf6k-4
  27. www.youtube.com/watch?v=XeWUtjzDIHw
  28. www.youtube.com/watch?v=CmHj--QJ8U4

Open-source projects that support Veo

calesthio

OpenMontage

Video studio that your AI coding assistant runs

Video framework
  • Google Veo
  • Kling
  • Seedance 2.0
  • +7
62k GitHub stars25 days agoAGPL-3.0

HKUDS

ViMax

Multi-shot AI video from an idea, a script, or a novel chapter

Video framework
  • Seedance 2.0 Fast
  • GPT Image 2
  • Google Omni
  • +4
13k GitHub starsyesterdayMIT

krillinai

OpenCreator

Dubs, translates and writes short-video scripts on your own machine

Video framework
  • Seedance 2.5
  • Kling v2.1 Master
  • Veo 3.1
  • +7
13k GitHub starstodayApache-2.0

Video diffusion papers and projects, sorted into labeled sections

Datasets & papers
  • Wan 2.1
  • HunyuanVideo
  • LTX-Video
  • +7
5.8k GitHub stars9 days agoNo license

SegFault42

HeliosGen

AI image and video pipelines on a local node canvas

ComfyUI workflow
  • Veo 3.1
  • Kling 3.0
  • Seedance 2.0
  • +7
2.3k GitHub stars14 days agoNo license

LingyiChen-AI

AIComicBuilder

Writes an animated comic drama from your script, shot by shot

Platform
  • OpenAI
  • Gemini
  • Kling
  • +4
1.9k GitHub stars5 months agoApache-2.0