Skip to main content
Gemini Omni Flash 1.1 Prompting Guide

Gemini Omni Flash 1.1.
Video with native audio.

Google Gemini Omni Flash 1.1 turns text, a start and end frame, reference images and clips, or an existing video into a short clip with native audio baked in, no separate sound step, at 360p draft up to 4K. It does text-to-video, first-to-last-frame, reference-to-video, video editing, and one-tap extension of a clip you already made. This guide covers how to prompt the picture and the sound together.

Overview

Omni Flash is Google's all-in-one short-form video model. The headline feature is native audio: ambient sound, foley, and atmosphere are generated together with the picture, so a single prompt produces a finished clip you can post. You write the scene and the sound in the same sentence.

It works five ways from one composer: text-to-video, frames (a start image, optionally with an end image the model interpolates toward), references (up to 7 stills and 3 clips for subject, style, or motion), video-to-video editing (hand it a clip and describe the change), and Extend, which grows any clip by 3 to 10 seconds from a source up to 30 seconds. Output is 16:9 or 9:16 at 360p, 720p, 1080p, or 4K, 3 to 10 seconds, billed per second of output: 1, 3, 4.5, or 9 credits by tier (1.25, 3.5, 5.25, or 10.5 when you edit or extend a clip).

Native

Audio generated with the video

3–10s

360p to 4K · 16:9 or 9:16

5

T2V · frames · references · edit · extend

Key Features

Picture and sound together

Native Audio, Always On

Omni Flash generates sound with the video — rainfall, footsteps, ambient hum, a quiet music bed. There's no audio toggle and no post step. Name the sounds you want in the prompt and they get rendered alongside the action. Generic 'with audio' does little; specific cues ('soft rainfall and distant city hum') land.

Wind, light, weather

Atmospheric Motion

Strong on naturalistic, atmospheric motion — breath in cold air, snow underfoot, light shifting through trees. The model reads sensory language as motion direction, not just visual style. Pair a clear subject + action with one dominant camera move for clean, coherent clips.

Animate a still, or interpolate between two

Start and End Frames

Drop in a start image and Omni Flash pins it as the first frame, keeping layout, palette, and identity while adding motion and sound. Add an end image and the model interpolates the motion between the two stills: orbits, reveals, loops. Each image under 4.77MB. Great for bringing a product shot, a scene, or a character frame to life.

Change a clip in place, or make it longer

Edit & Extend

Hand it an existing clip (3–10s) and describe the change for a video-to-video edit that keeps the source length, or hit Extend on any result to add 3–10 seconds of what happens next. Extension reads up to 10 seconds of the source for continuity and accepts sources up to 30 seconds, so a shot can grow across several passes.

Example Videos

Each example shows the exact prompt that produced the result. Copy any prompt with one click.

Text → Video with Audio

720p · 16:9 · 8s

A neon-lit Tokyo side street in the rain at night, reflections shimmering on wet asphalt, a person under a clear umbrella walks slowly past glowing signage, soft rainfall and distant city hum, gentle cinematic push-in, photoreal.

Lead with the scene, then name the sound ('soft rainfall and distant city hum') in the same breath as the visuals — Omni Flash renders both. One camera move ('gentle cinematic push-in') keeps the motion clean. 'Photoreal' anchors the look.

Atmospheric Nature

720p · 16:9 · 8s

A red fox trots through a snowy pine forest at golden hour, breath visible in the cold air, snow crunching underfoot and soft birdsong, low tracking shot following the fox, warm backlight through the trees.

Sensory detail doubles as motion and sound direction: 'breath visible', 'snow crunching underfoot', 'soft birdsong'. The named audio cues come through. 'Low tracking shot following the fox' gives the model one clear camera instruction to execute.

Vertical Social Clip

720p · 9:16 · 8s

Close-up of a barista pouring latte art into a ceramic cup in a cozy cafe, steam curling upward, the soft hiss of the espresso machine and quiet acoustic music, shallow depth of field, warm morning light.

9:16 for Reels / TikTok / Shorts. Close-up + shallow depth of field reads well at 720p. The audio prompt layers two distinct sounds (machine hiss + acoustic music bed) — Omni Flash mixes them rather than picking one.

Image → Video

720p · 16:9 · 8s · start frame

Bring this desk scene to life: code scrolls subtly on the laptop screen, steam drifts up from the coffee mug, soft keyboard clicks and a quiet room ambience, gentle slow push-in, keep the layout and lighting consistent with the reference.

With a start frame, describe the MOTION you want added, not the whole scene — the image already supplies the composition. 'Keep the layout and lighting consistent with the reference' holds it steady. Subtle, specific motions ('code scrolls', 'steam drifts') beat big ones. Add an end frame when you know where the shot should land.

Prompting Tips

Write the sound into the prompt

Audio is native, so treat it as part of the scene. Name specific sounds — 'distant city hum', 'snow crunching underfoot', 'soft espresso machine hiss' — rather than a vague 'with sound'. Specific cues get rendered; generic ones get ignored.

One subject, one camera move

Lead with subject + action, then name a single dominant camera move ('gentle push-in', 'low tracking shot following the fox'). Stacking two motion types into a 3–10s clip usually produces compromised, jittery motion.

For frames, prompt the motion

The start frame already defines the composition — your prompt should describe what MOVES and what you HEAR, plus a line to keep it consistent with the reference. With an end frame, describe the path between the two. Frames and references must each be under 4.77MB; compress large uploads first.

Draft at 360p, finish at 1080p or 4K

360p is a third of the price of 720p and much faster. Use it to test a prompt, then re-run the keeper at 720p, 1080p, or 4K. The higher tiers are upscaled by the model, so nail the composition before paying for pixels.

Pick aspect by channel

16:9 for landscape / YouTube. 9:16 for vertical social. Choose the aspect ratio up front and the model composes the framing accordingly; when you edit or extend a clip, the output follows the source.

Keep clips short and specific

3–10 seconds is the whole range. Shorter, tightly-described clips are more reliable than long ones trying to cram in multiple beats. For a longer sequence, generate several focused clips and stitch them.

Settings Reference

SettingValuesNotes
ModesText-to-video · Frames · References · Edit · ExtendOne composer. Frames, references, and a source clip are mutually exclusive; reference images and clips can be combined.
AudioNative, always onGenerated with the video. No toggle. Prompt the specific sounds you want.
Duration3–10 secondsFor Extend, the seconds added; an in-place edit keeps the source length. Billed per second of output.
Resolution360p · 720p · 1080p · 4K1, 3, 4.5, or 9 credits per second; 1.25, 3.5, 5.25, or 10.5 when editing or extending a clip. 1080p and 4K are upscaled by the model.
Aspect ratio16:9 · 9:16Choose up front; the model frames to the ratio.
ReferencesUp to 7 images and 3 clips, images under 4.77MBNot alongside frames or a source clip. Compress large uploads.
Extend source1–30 secondsAdds 3–10s per pass; the whole output is billed.

FAQ

Yes — audio is native and always on. Ambient sound, foley, and a light music bed are generated together with the picture from your prompt. There's no separate audio step and no toggle. Name the sounds you want and they get rendered into the clip.