Use case · AI UGC video

AI UGC video ads from two photos

A UGC ad is someone talking to a phone camera with your product in their hand. Here it is generated instead of filmed: you bring a product photo and a person, the model builds the opening frame, then animates it with the voice and the room sound in the same pass. A clip runs $0.09–$1.00 on Gemini Omni Flash and up to $3.00 on Veo 3.1 with sound, the frame costs $0.11, and none of it is a subscription.

Made for this page, start to finish: two reference photos → a Nano Banana Pro opening frame → eight seconds of Gemini Omni Flash, speech and room tone generated with the picture. $1.13 for the take you are watching. Play it with sound on.
01What it is

What is an AI UGC video?

An AI UGC video is a short ad in user-generated-content format — a person speaking to a phone camera, holding a product, in a kitchen or living room instead of a studio — generated by a video model rather than shot with a creator. The production is three calls: reference photos of the product and the person, one still frame that locks who is on screen and what they are holding, then a video model that animates that frame with synchronized speech. Two models cover it here. Gemini Omni Flash is the cheap one and the newer one — Google's model card puts real-world physics simulation among its abilities and complex motion among its remaining weaknesses, which matches what I see: ordinary handling usually survives, elaborate handling is a coin flip. Veo 3.1 is what you reach for when a shot depends on an object behaving. One clip is $0.09–$1.00 on Omni Flash, $1.00 on Veo 3.1 Fast and $3.00 on the full Veo 3.1 for eight seconds with sound, up to $6.00 at 4K — and it renders in a couple of minutes instead of the two to four weeks a creator brief usually eats.

What does one clip actually cost?

The bill for the clip above, unrounded: $0.11 for the product reference, $0.11 for the person, $0.11 for the opening frame that puts them together, $0.80 for eight seconds of Omni Flash — $1.13 in total. The vertical six-second cut further down cost $0.60. Re-recording the voice separately runs $0.01 per 200 characters of script, about a cent a sentence. New accounts get $0.20 free, which covers a reference photo and a first frame — enough to see whether the character works before you pay for video.

That is the first-take number, and first takes are not a plan. Some clips come out on the first attempt; plenty need two or three, because a line lands flat, a hand does something strange or the model refuses the wardrobe. Then there is the other thing nobody puts in the price: an ad is rarely one clip. A 20–30 second spot is usually three or four takes cut together — hook, product in hand, close-up, sign-off — so the realistic bill for a finished ad is a few dollars, not a few cents, and it goes up fast if you shoot the motion-heavy shots on the full Veo 3.1 at $3.00 each. Cutting them together is the free part: ffmpeg trims, joins and lays audio over video from the command line, no editor involved.

02The workflow

Two photos in, one ad out

  1. Bring real references

    A clean product shot and a photo of the person. Anything soft, cropped or oddly lit in the input comes back multiplied in the video.

  2. Build the opening frame

    Generate one still holding both references — the person with the product, framed the way the ad should open. $0.11 on Nano Banana Pro, and you can redo it until it looks right.

  3. Animate that frame

    Send the still in as the first frame and describe the motion and the line she says. Omni Flash charges $0.10 per second of 720p with audio; Veo 3.1 costs more and is the one to try when a single take has to hold several actions in a row.

AI-generated product reference photo: an unbranded amber glass serum dropper bottle with a cream ceramic cap on pale linen
Reference 1 — the product, one clean angle, even light. Nano Banana Pro, $0.11.
AI-generated creator reference portrait: a woman in a cream knit sweater sitting in a bright living room
Reference 2 — the person. Same light and same wardrobe you want in the ad.
AI-generated UGC opening frame: the woman from the reference portrait holding the amber serum bottle beside her shoulder, phone-camera framing
The opening frame, generated from both references at once — this still is what the video model continues.
03References

Everything downstream is only as good as what you feed in

This is the part people skip, and it is the one that decides the result. A model given a blurry three-quarter phone snap of a bottle does not fail loudly — it quietly invents the side it cannot see, and that invented side is what ends up on screen for eight seconds. Give it a flat, evenly lit product shot with the label facing camera, and the same bottle survives the frame step, the video step and the edit after it.

The person matters the same way. Google's Veo documentation is blunt about it: choose clear, well-lit images that show the subject from the angle you want, and keep the lens and lighting language identical between shots. My rule of thumb after a few dozen of these: if the reference portrait and the intended scene disagree about where the light comes from, the face will drift somewhere in the middle and stop looking like the same person.

One good product angle

A single flat, evenly lit shot beats five casual snaps. Glare, motion blur and a hand covering half the label are all reconstructed as invention.

Same person, same look

Keep the haircut, the wardrobe and the makeup consistent across reference photos. Mixed looks average into a face that matches neither.

Match the light

A flash-lit portrait dropped into a soft-daylight kitchen scene reads as a composite. Pick the light in the reference that the ad will actually have.

Watch small type

Nano Banana Pro holds short label text; a dense ingredient list turns into texture. If the type matters, plan a close-up shot with real packaging instead.

Framing comes later

References do not need to be composed like an ad. Composition is the job of the opening frame — that is the step where you argue with the model.

Don't hand it a crop

A 300-pixel crop pulled off a marketplace listing gives the model almost nothing to keep. Full-resolution originals, every time.

04First frame

Why the opening frame beats raw references

Both routes work: you can hand the video model your reference photos and let it compose the scene, or you can generate one still first and animate that. In practice the second is better, and the reason is mundane — a video model asked to invent the composition invents it every frame, while a model handed a still only has to continue it. On Gemini Omni Flash the gap is large enough that I stopped using references-only for product shots.

Two clips below, same prompt, same references, six seconds each, $0.60 apiece. The only difference is whether an approved still went in as the first frame. Nothing else was changed and neither clip was re-rolled.

References only

No opening frame. The cream cap disappears within the first second, the bottle turns into a different dark bottle, and by the end it has drifted out of the shot. Her face is fine — the light on it is arguably better than in the take next door. The product is what breaks.

From an approved first frame

Same prompt, started from a still we looked at first. Cap, glass, framing and camera distance hold for the full six seconds, because the model was continuing a frame instead of guessing one.

The honest version: references-only is not broken — the face held up fine, and for a scene without a product in hand it is often enough. If anything the light is nicer in that take: softer on her face, less window contrast, because the model was composing a room it liked instead of continuing ours. One caveat worth stating, since it changes the reading — our creator is generated, not a photograph of a real person, so the model had room to relight her. A real portrait pins the light and the skin much harder, and the references-only route has less freedom to flatter it. It is control you are buying here, not quality. On Veo 3.1 the choice is made for you: Google's API takes either a first frame or up to three reference images, never both in the same request, and reference images require the eight-second length.

05Model choice

Match the shot to the model

Omni Flash is very good at the format UGC actually lives in: a person talking to the lens. Interviews, podcast-style pieces, a founder explaining a product, someone holding a jar and describing it — and it renders the mouth, the micro-movement and the room tone in one pass, which makes it the cheapest way to get a synced voice out of a video model. It is also the newer model of the two, and Google sells it on physics: the announcement talks about an intuitive grasp of gravity, kinetic energy and fluid dynamics, and the model card lists real-world physics simulation as an ability while naming complex motion as a known weak spot. Both halves of that are true in practice. Things fall, liquids pour, weight reads correctly — right up until the shot asks for a chain of precise actions with an object.

So the split is not talking heads versus motion; it is how much choreography a single take has to hold. The test below is the hard end of that scale, not a verdict on either model: same opening frame, same prompt, one model each — uncap the bottle, squeeze the pipette, two drops into the open palm, all in eight seconds.

Gemini Omni Flash · 8 s · $0.80

The uncapping is convincing, the pipette comes out cleanly and the drops do land in the palm — the fluid part is fine. The object is not: the bottle vanishes from her lower hand the moment the palm opens, comes back for the recap, and the clip ends with her empty-handed. One of the four actions dropped out.

Veo 3.1 Fast · 8 s with audio · $1.00

Identical prompt and identical first frame, and note this is Veo 3.1 Fast — the cheap tier, not the full model. It keeps the bottle through all four actions. There is a half-second where a third hand ghosts through the frame, so it is not artifact-free either; it just held the sequence together.

Talking to camera → Omni Flash

Lip movement, room tone and speech arrive together, $0.10 per second. For a review, a testimonial or a two-person exchange it is the obvious pick.

Long action chains → try Veo too

One motion Omni handles; four in a row is where takes start dropping actions. Our comparison used Veo 3.1 Fast at $1.00 for eight seconds with audio, the full Veo 3.1 is $3.00 — generate both and keep the one that works.

Wardrobe changes the verdict

Omni's safety filter is noticeably stricter about clothing: a fitness creator in a sports bra gets refused there far more often than on Veo 3.1, same prompt, same reference. Plan the wardrobe or plan the model.

Length and resolution

Omni Flash is 3–10 seconds per clip, exactly as many as you pay for — 720p on the first generation, 360p to 4K on Omni 1.1, which can also extend a scene to 40 seconds. Veo 3.1 does 4, 6 or 8 seconds up to native 4K, and its clips can be extended past the first cut.

Sound: always vs optional

Every Omni clip has audio whether you asked or not. On Veo audio is a toggle — silent renders are cheaper, which matters when the voice is coming from elsewhere anyway.

Refusals cost nothing

A blocked generation is refunded automatically on both models. The expensive part of a refusal is the four minutes, not the dollar.

06Voice & edit

When the voice needs a second take

Omni Flash writes the speech into the same pass as the picture, and for conversational English it is usually fine. Brand names are where it wobbles — invented product names, foreign words, model numbers, anything a native speaker would have to be told how to say. The bigger version of this problem is languages other than English, and every video model has it: Russian-speaking reviewers testing Seedance 2.5 describe a line that starts acceptable and comes apart on the way — unstressed vowels reduced wrong, stress landing on the wrong syllable, and by the end of a fifteen-second scene an accent that is closer to a Slavic hybrid than to Russian, because Russian, Ukrainian and Belarusian sit close enough for the model to blend them. Some words come out clean, the next one gives the whole take away. If your script is not in English, listen to the whole thing before you sign it off.

The fix is to stop asking the video model for the voice and generate it separately. Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview) runs on the same balance through our MCP server: $0.01 per 200 characters of script, 30 voices, one speaker or a two-person dialogue, 24 kHz WAV out. Direction is plain English — persona, pace, accent, and the phonetic spelling of any word you need said correctly. Then drop the track over the clip in any editor. Keep your rewritten line about as long as the original take and the mouth still roughly agrees; where it doesn't, cut to the product.

Generated voice-over · $0.01

A 200-character script read by Gemini 3.1 Flash TTS with one direction line: friendly young woman, phone review in her kitchen, slightly fast, and pronounce the invented brand name the French way. Thirteen seconds of audio, one cent, no video model involved.

Putting the ad together needs no editor either. ffmpeg joins takes end to end, trims the half-second where a hand does something strange, swaps the generated audio for your TTS track, drops in a crossfade and exports the 9:16 version — all from one command line, free, on any machine. The syntax is famously unfriendly, which is exactly why it pairs well with a model: describe the cut you want to Claude or any LLM, get the invocation back, run it. For a three-clip UGC spot that is the entire post-production stage, and it is the reason the numbers on this page stay at the generation cost instead of adding a subscription on top.

07Pricing

What an AI UGC ad costs here

Per-unit prices for the models used in an AI UGC video workflow
ModelPriceUnitAudio
Nano Banana Proreferences & first frame$0.111K–2K image
Nano Banana 2drafts and variants$0.03–$0.13512px–4K image
Gemini Omni Flash$0.10 per second$0.30–$1.003–10 s, 720palways on
Veo 3.1 Fastwith audio$0.50–$1.004–8 s, 720p/1080poptional
Veo 3.1 Litecheapest video$0.10–$0.364–8 s, 720poptional
Gemini 3.1 Flash TTSMCP only$0.01per 200 charactersvoice only

Veo prices are for 720p and 1080p; 4K costs more. Omni Flash bills $0.10 for every second of output, so the clip length is the price. Full tables on the pricing page.

Top-ups are crypto from $1, or card, PayPal and SEPA from $20. Deposits of $50 or more add 5%, $100 or more adds 10%, and an active promo code adds another 10% of the same deposit — so $100 with a code lands as $120 on the balance. At $1.13 for an eight-second clip that is around a hundred takes — call it twenty-five to thirty finished ads once retries and multi-clip cuts are counted. The +5% / +10% volume bonus applies to crypto top-ups only.

Every model and resolution is listed on the full pricing page, and the Gemini Omni Flash page covers what that model can and cannot do.

08Learn more

Guides from the blog

09FAQ

Questions people actually ask

Do I need a real person to make an AI UGC video?

No. The creator can be generated too — every person you see on this page came out of a text prompt and never existed. If you do use a real person's face, you need their permission: likeness rules apply to generated video the same way they apply to a photoshoot, and platform policies on synthetic likeness are stricter than most people assume.

Will the same face come back for the next ad?

Yes, as long as you keep the same reference photos and reuse the same opening frame. Regenerating the frame from the same references gets you close but not pixel-identical — the safest habit is to keep every approved first frame and animate it again rather than rebuild it.

Do I have to disclose that the ad is AI-generated?

On TikTok, yes: its advertising policy on edited media and AI-generated content requires either the AIGC label from Ads Manager or your own visible disclaimer on the creative. Meta applies its own AI labels automatically to ads it detects as generated. Both policies changed during 2026, so check the current version before a large flight rather than trusting a blog post — including this one.

Which model should I use for a talking-head UGC ad?

Start with Gemini Omni Flash at $0.10 per second, audio included — for a person speaking to camera it is both the cheapest and the least fussy, and it is the newer model, with physics Google explicitly advertises. Do not read that as "Omni can't do motion": it pours, it drops, it carries weight. What it drops is long chains of precise handling in one take — on our four-action dropper test it lost the bottle while Veo 3.1 Fast kept it. So for a hook or a testimonial, Omni. For a shot built on choreography, generate it on both and keep the take that works; the failed one costs cents.

Can I use footage I filmed myself?

Yes, up to 10 seconds of it. Upload your clip and Omni Flash will edit it as video-to-video — change the lighting, the background, what the person is holding — while keeping the take. Longer sources get trimmed to the model's limit before generation.

What happens if a generation gets refused?

The balance is refunded automatically, on both Omni Flash and Veo. Refusals cluster around clothing, minors and recognisable real people; the wardrobe rule from the model-choice section above is the one that saves the most retries.

Read the full answer
Can I run this from Claude, Cursor or my own script?

Yes — the same models are exposed over our MCP server, including the speech model, which has no web interface at all. Reference images and first frames can be passed as job IDs, public URLs or inline base64, so a whole UGC batch can be scripted from an agent.

What does a whole 30-second ad cost, not one clip?

Count takes, not clips. A 20–30 second spot is normally three or four clips cut together, and each one takes a couple of attempts before it is usable — a flat line reading, a hand doing something odd, a refusal. At $0.80 for eight seconds of Omni Flash that puts a finished ad in the low single digits; shooting the motion-heavy cuts on the full Veo 3.1 at $3.00 each moves it into the tens. The editing adds nothing: ffmpeg joins the takes, trims them and lays the voice-over on top from the command line, and an LLM can write those commands for you.

Two photos and a couple of dollars

Sign up free with $0.20 on the balance, top up with crypto from $1 or a card from $20. First frame in about a minute, first clip in five.