Model · Google Gemini Omni Flash

Gemini Omni Flash — AI video with sound, from $0.30 a clip

Google's world model for short-form video: every clip arrives with synchronized audio, physics-aware motion, and editing by conversation — describe a change and the scene regenerates. You set the length exactly: 3 to 10 seconds per clip, 16:9 or vertical 9:16 for Shorts and Reels; the 1.1 generation adds 360p drafts, 1080p and 4K, and scene extension to 40 seconds. $0.10 per second at 720p ($0.30–$1.00 a clip), paid with crypto or card. No subscription.

An actual Gemini Omni Flash clip generated on BananaBanana — one text prompt, a full 10 seconds, 720p, audio generated in the same pass. At $0.10 a second that is $1.00; a 3-second draft would be $0.30. Press play with sound on.

Omni 1.1 Flash is live

The second generation adds 360p drafts at a third of the price, 1080p and 4K output, a closing frame, video references and scene extension to 40 seconds. It has been the model behind Omni Flash in the studio since September 2026.

Read the Omni 1.1 Flash page
01What it is

What is Gemini Omni Flash?

Gemini Omni Flash is the first model in Google's Gemini Omni family — a world model that reasons about a scene the way a physics engine does, then renders it as video. You send text, an image, or a previous result; it returns a 720p clip at 24 fps with native synchronized audio, exactly as long as you asked (3 to 10 seconds), — ambience, effects, even speech — generated in the same pass as the pixels, not added afterwards. Google shipped it to the Gemini app, Flow and YouTube Shorts in May 2026 and opened the API (gemini-omni-flash-preview) on June 30, 2026. On BananaBanana it runs in production next to the Veo 3.1 family: same balance, $0.10 per second of video ($0.30–$1.00 per clip), paid with crypto or card.

Why is it built for Shorts, Reels and TikTok?

Short-form video needs three things: sound on by default, believable motion, and fast iteration. Omni Flash delivers all three. Audio is always on — the model decides what the scene should sound like from what it shows. Motion comes from a world model that reasons about the scene: syrup pours, paper drifts, steam rises where it should — usually without prompt gymnastics (for the most physics-critical shots, Veo 3.1 is still the safer render). And instead of re-prompting from scratch, you refine by conversation — “make it dusk”, “add rain on the window” — while the model keeps the character, lighting and camera consistent. The prompt is also the control panel: left alone the model likes to cut between shots, so you ask for a single unbroken take, mark the beats as [0-3s] ranges, and quote any text that should appear on screen. Generate in vertical 9:16 and the clip drops straight into Shorts, Reels or TikTok with no crop.

One Omni Flash generation in vertical 9:16 — pad thai tossed in a wok over a street gas burner, steam rising, with the sizzle and market ambience generated in the same pass. $1.00 for the full 10 seconds, ready for Shorts, Reels or TikTok with no crop.
02Capabilities

What Omni Flash can do

Always-on synchronized audio

Every clip ships with generated sound — ambience, effects, speech. Steer it with an "Audio:" line in the prompt.

Conversational editing

Describe a change and the scene regenerates with character, lighting and camera intact. Edits chain turn after turn — and any finished video, even a Veo clip or your own upload, can be sent back in for a video-to-video edit.

Physics-aware world model

The model reasons about the scene like a physics engine — liquids, smoke and cloth usually behave plausibly on the first try.

Image-to-video

Any photo becomes the opening frame — your own generation or an upload — while the prompt describes the motion and the soundscape. It can be combined with reference images in the same request.

Subject references

Up to 10 reference images keep characters or products consistent in a completely new scene — and they can be combined with a starting frame.

You set format and length

Pick the aspect ratio per generation — vertical output is framed for Shorts, Reels and TikTok — and the exact clip length, 3 to 10 whole seconds, paying $0.10 for each of them.

Another single Omni Flash generation — syrup pouring onto pancakes, sound included (turn it on: the drizzle is generated too). One prompt, no audio post-production, $1.00 for 10 seconds at $0.10 a second.
03Limitations

What it can't do — honestly

These are limits of Google's API itself, not of our integration — and they matter before you spend a dollar. 720p is the native render: the first generation stops there, and the 1080p and 4K tiers of Omni 1.1 are upscales of the same 720p pass, not extra detail — for native 4K use Veo 3.1. A single clip is 3–10 seconds: the first generation cannot go past that at all, and Omni 1.1 only gets further by appending 3–10 second extensions, up to 40 seconds in total. Editing changes a scene, it never lengthens it: an edited clip always comes back exactly as long as its source.

Also missing from the preview API: seeds (the same seed and prompt still give a different take), negative-prompt parameters, prompt enhancement, multiple samples per request, and switching the audio off. If any of those are dealbreakers, the Veo 3.1 family covers them — same balance, same generator.

04How it works

From prompt to clip in minutes

  1. Sign up free

    Email and password — no credit card, no Google account, $0.20 free starting balance.

  2. Top up with crypto or card

    From $1 in USDT, USDC, BTC, ETH, SOL, TON and more. +5% bonus from $50, +10% from $100.

  3. Generate, then talk to it

    Write the prompt (add an "Audio:" line), pick 16:9 or 9:16 and the exact length, generate from $0.30 — then refine by describing changes.

05Pricing

Priced by the second, audio included

Gemini Omni Flash price compared with Veo 3.1 clips with audio
ModelPrice per clipDurationResolutionAudio
Gemini Omni Flash$0.10 per second$0.30–$1.003–10 s (you pick)720palways on
Veo 3.1 Fastwith audio$0.50–$1.004–8 s (you pick)up to 4Koptional
Veo 3.1with audio$1.50–$3.004–8 s (you pick)up to 4Koptional

Veo prices shown for 720p/1080p; 4K costs more. Omni Flash is billed $0.10 per second of output, so a clip costs $0.30–$1.00; an edit is a new generation of the same length as its source, at the same rate. Full tables on the pricing page.

Compare every model on the full pricing page — or read the hands-on Omni Flash API review before you spend.

06Learn more

Guides from the blog

07FAQ

Gemini Omni Flash, answered

What is Gemini Omni Flash?

Gemini Omni Flash is the first model in Google's Gemini Omni family — a world model that turns text, images or a previous clip into a short 720p video with synchronized sound generated in the same pass. It shipped to consumers in May 2026 and reached the API as gemini-omni-flash-preview on June 30, 2026. On BananaBanana it runs in production alongside the Veo 3.1 family.

How much does a Gemini Omni Flash video cost?

$0.10 per second of video on BananaBanana — $0.30 for a 3-second clip, $1.00 for the full 10 seconds, audio included, no subscription. An edit of an existing clip is a new generation of the same length, at the same rate. You top up with crypto from $1, or by card from $20, and pay per clip.

Read the full answer
Does Gemini Omni Flash generate sound?

Yes — always. Every clip ships with native synchronized audio: ambience, effects, even speech if you ask. You cannot turn it off; the only control is text, so describe the soundscape in an "Audio:" line inside the prompt.

Can I choose the clip length or resolution?

The length, yes: any whole number of seconds from 3 to 10, and you pay $0.10 for each of them at 720p. Resolution depends on the generation: Omni Flash 1.0 renders 720p only, while Omni 1.1 Flash adds a 360p draft tier at a third of the price and upscaled 1080p and 4K at a higher per-second rate. You also choose the aspect ratio: 16:9 for widescreen or 9:16 vertical for Shorts, Reels and TikTok. One exception: an edited clip inherits the length of its source, because the model cannot stretch or shorten a video.

Read the full answer
How does conversational editing work?

Generate a clip, then describe the change — "make it snow", "swap the mug for a glass" — and Omni Flash regenerates the scene while keeping character, lighting and camera consistent. Edits chain on top of each other. You can also edit a video it didn't make: send in any finished clip, including a Veo generation or your own upload, and it comes back restyled at the same length. Each edit is a normal generation at $0.10 per second. Making a clip longer is a separate feature — scene extension — and it exists only in Omni 1.1 Flash.

Should I use Gemini Omni Flash or Veo 3.1?

Pick Omni Flash when sound matters and the format is short vertical or social video: audio is always on, physics look right, and iteration happens by conversation at $0.10 a second ($0.30–$1.00 a clip); pick the 1.1 generation for 360p drafts, 4K delivery or scenes longer than ten seconds. Pick Veo 3.1 when you need control — exact 4/6/7/8-second durations, native 4K, optional audio, seeds, first/last frames and extensions to 148 seconds.

Do I need a Google account or credit card?

No. Sign up with an email and top up with crypto from $1 — USDT, USDC, BTC, ETH, SOL, TON and other major coins — or pay by card, PayPal or SEPA from $20. Then generate; unused balance never expires.

Make your first clip with sound

Gemini Omni Flash — $0.30–$1.00 per clip with always-on audio, billed $0.10 a second; Veo 3.1 from $0.10. Pay as you go with crypto or card, no subscription, free $0.20 to start.