Skip to main content
Flux 3 Video Prompting Guide

Flux 3 Video.
Picture and sound in one pass.

Flux 3 Video is the video model from Black Forest Labs, and it writes the soundtrack while it writes the picture. Start from a prompt, from a still frame, from two stills, or from a clip you already have. Up to 20 seconds, six aspect ratios including 21:9 ultrawide, 720p or 1080p.

Overview

Most video models hand you a silent clip and leave the sound to you. Flux 3 Video generates audio with the picture, in the same pass, matched to what is on screen. Footsteps land on footfalls, a door closes when the door closes, and room tone sits under the whole shot. You direct it from the prompt, so sound is something you write rather than something you add later.

There are three ways in, all on the same page. Write a prompt and you get text-to-video. Attach a start frame and you get image-to-video, with an optional end frame the model interpolates toward. Attach a clip instead and the model continues or transforms it. Clips run 5 to 20 seconds at 720p (5 credits per second) or 1080p (8 credits per second). Video extension is capped at 15 seconds and runs at a higher rate.

20s

Longest clip

6

Aspect ratios

Audio

Generated with the picture

Key Features

Sound written with the picture

Native Synchronized Audio

Audio is generated in the same pass as the video, so it lines up with what happens on screen instead of floating over it. Ambience, effects, and voices are all prompt-directed. Turn it off with one toggle when you plan to score the clip yourself.

Text, frames, or an existing clip

Three Entry Points

Leave the image and video slots empty and your prompt is a text-to-video job. Attach a start frame and it animates that frame. Attach a clip and the model continues it. One page, one prompt box, no mode dropdown to get wrong.

Interpolate between two stills

Start and End Frame

Give a start frame and an end frame and the model builds the motion that gets from one to the other. Useful for a controlled reveal, a match cut between two renders, or a transformation where both ends need to be exact.

The longest in the Flux family

Up to 20 Seconds

A single clip runs 5 to 20 seconds. That is room for a real beat structure, a full product spot, or an atmospheric shot that gets to breathe. Video extension tops out at 15 seconds per job, and you can chain extensions from there.

21:9 ultrawide through 9:16

Six Aspect Ratios

21:9 for anamorphic and title sequences, 16:9 for standard landscape, 4:3 and 3:4 for classical proportions, 1:1 for square social, 9:16 for vertical. The ultrawide option is rare among video models and it is native, not a crop.

5 or 8 credits per second

720p and 1080p

Draft at 720p, finish at 1080p. The higher tier holds fine texture that matters on atmospheric work, product materials, and anything with small on-screen detail. For character and action work, 720p usually carries the shot.

Example Videos

Each example shows the exact prompt that produced the result. Copy any prompt with one click.

Cinematic Drone Shot

21:9 ultrawide · audio on

Ultra-wide cinematic drone shot sweeping across a vast desert canyon at golden hour, anamorphic look, dust catching the light, deep engine-less silence broken by wind

One sweeping move across a wide subject is what 21:9 is for. Note the audio direction at the end: naming the silence, then the one sound allowed to break it, gives a cleaner bed than asking for wind alone.

Character Close-Up

16:9 · audio on

Close-up of an elderly fisherman's weathered face as he looks out to sea at dawn, gentle camera drift, seagulls calling in the distance

Close-ups want a gentle move and a specific face. "Weathered" and "elderly fisherman" do more work than a paragraph of description, and the distant gulls put the shot somewhere without adding anything to the frame.

Stylized Anime Scene

16:9 · audio on

Anime-style schoolgirl running along a cherry-blossom lined river path, petals swirling in the wind, vibrant spring colors, upbeat mood

Style prompts still need a motion verb. "Running" plus "petals swirling" gives the model two moving layers, which is what separates an animated shot from a pan across a still illustration.

Product Spot

1:1 square · audio on

Square product spot of a luxury wristwatch rotating on a black pedestal, dramatic rim lighting sweep, dust particles in the beam, deep ambient tone

Let the subject move instead of the camera when the product is the whole shot. The rotating pedestal plus a lighting sweep covers the object from every angle, and dust in the beam makes the light readable.

Atmospheric Landscape

16:9 · audio on

Northern lights dancing over a snow-covered pine forest, stars visible, slow upward tilt, ethereal ambient tones

Atmospheric work runs on one slow move and one thing that animates on its own. The upward tilt reveals the sky while the aurora does the moving, so nothing in the frame is waiting around.

Vertical Action

9:16 vertical · audio on

Vertical shot of a parkour athlete vaulting across rooftop gaps at sunset, dynamic tracking camera, city sprawling below, sneakers scraping concrete

For 9:16 action, name the tracking camera and one contact sound. "Sneakers scraping concrete" is a specific event tied to the movement, which reads far better in the mix than a generic city ambience.

Prompting Tips

Write the sound, do not leave it to chance

Flux 3 Video generates audio whether or not you describe it. If you say nothing, you get its guess. Name one to three specific sounds at the end of the prompt and you get the bed you wanted instead.

Say "no music" when you want ambience only

Scores show up uninvited on cinematic prompts. Ending the prompt with "no music" keeps it to ambience and effects, which is what you want if the clip is going into an edit with its own track.

One camera move per clip

Pick one: slow push, lateral drift, crane up, 90 degree orbit, handheld follow. Stacking two moves in one prompt produces confused motion. If you need a second move, that is a second clip or an extension.

Use motion verbs, not adjectives

The prompt should read like stage directions. "She exhales, fogs the glass, then wipes a circle clear" plays. "A contemplative woman on a bus" gives the model nothing to animate and you get a near-still frame.

Order your beats for the duration you picked

A 20 second clip needs three or four beats or it drifts. A 5 second clip needs one. Match the number of actions to the length rather than paying for seconds the prompt does not fill.

Draft at 720p, finish at 1080p

At 5 credits per second, 720p is where you find the shot. Once the motion and audio are right, rerun the same prompt at 1080p for the version you ship. Same prompt, same structure, more detail.

Aspect ratio only applies to text-to-video

Attach a start frame or a source clip and the ratio comes from that asset. If you need 21:9 from an image, crop the image to 21:9 before you upload it.

Turn audio off if you are scoring the clip

For montage work where every clip gets the same track in post, switch audio off. There is no reason to generate a bed you are going to mute.

The Five-Part Prompt Framework

1. Scene and subject. Open with what we are looking at and how close we are. "Medium close-up of a woman at a rain-streaked bus window at night" fixes the framing, the subject, the setting, and the time of day in one line. Vague openings produce generic footage.

2. Camera movement. Name exactly one move and its speed. "The camera holds one steady forward push", "one slow 90 degree orbit", "handheld follow, dipping with the ride". Speed matters as much as direction, and a clip with no named move tends to sit still.

3. Motion and action verbs. Write the beats in order, as stage directions. "She exhales, fogs the glass, then wipes a small circle clear." One beat for a 5 second clip, three or four for 15 to 20 seconds. This is the part that turns a moving photograph into a shot.

4. Audio cues. The model generates sound, so direct it. Close the prompt with the sounds you want, named specifically: "seagulls calling in the distance", "sneakers scraping concrete", "deep engine-less silence broken by wind". One to three cues is usually enough, and add "no music" when you want ambience only.

5. Style and mood. Close with the look: palette, lighting, lens, grade, rendering style. "Cool blue shadows with a warm sunrise edge, anamorphic cinematic grade" or "flat cel shading, visible line art, warm amber palette". Style last keeps it from crowding out the action.

Image to Video: Start Frame, End Frame, and Transitions

Attach a start frame and Flux 3 Video animates from it. The frame carries the subject, the composition, and the entire look, so the aspect ratio setting is ignored and the image defines the frame. Crop the image to the shape you want before uploading. Image-to-video is billed at the text-to-video rate and supports the full 5 to 20 second range.

Prompt the motion, not the picture. This is the single biggest mistake on image-to-video. The image already says what the scene looks like, so re-describing the subject only invites the model to drift away from it. Spend the whole prompt on what moves, how the camera behaves, and what it sounds like.

Add an end frame and the model builds the transition between the two stills. Both frames are honored, so the clip resolves exactly onto your closing image. Use it for a controlled reveal, a match cut between two renders, or a before-and-after transformation where both ends have to be exact. An end frame requires a start frame; it cannot be used on its own.

When you use both frames, describe the journey rather than either endpoint. "The camera pulls back as the fog clears and the light shifts from blue to amber" tells the model how to get there. Naming things that are already visible in either still is wasted prompt.

Video Extend: Continue or Transform a Clip

Attach a clip instead of an image and the job runs in video-to-video mode. The model picks up from your footage and writes a continuation, or transforms it if the prompt asks for a change of style or setting. The source clip defines the frame, so aspect ratio is ignored here too.

Extension is capped at 15 seconds per job and runs at a higher rate: 10 credits per second at 720p, 12 at 1080p. That premium buys you a continuation that matches the source rather than a fresh generation you have to cut around. For a longer sequence, extend the extension.

Prompt the continuation, not the recap. "The camera keeps pushing forward and the canyon opens onto a lit valley" gives it a direction to go. Describing what already happened in the source clip wastes the prompt and can make the model repeat the beat you just paid for.

Frame images and a source clip cannot be combined in one job. Pick one entry point. If you want to steer the end of an extension toward a specific still, generate the extension first, then run a separate start-and-end-frame job from its last frame.

Settings Reference

SettingValuesNotes
Promptstring (required)Scene, camera move, action beats, audio cues, style. See the five-part framework above.
Start frameimage URL (optional)Switches to image-to-video. Billed at the text-to-video rate. The image defines the frame, so aspect ratio is ignored.
End frameimage URL (optional)Requires a start frame. The model generates the transition between the two stills.
Source videovideo URL (optional)Switches to video extend. Higher credit rate, 15 second cap. Cannot be combined with frame images.
Resolution720p · 1080p5 or 8 credits per second on text and image jobs. 10 or 12 per second on video extend.
Aspect ratio21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16Text-to-video only. Ignored when a start frame or source clip is supplied.
Duration5 to 20 seconds5 to 15 seconds when a source video is supplied. Match your beat count to the length.
Generate audioOn / OffOn by default. Turn it off when the clip is going into an edit with its own track.

FAQ

Yes. Sound is produced in the same pass as the picture and is matched to what happens on screen, so effects land on the action rather than drifting. You direct it from the prompt by naming the sounds you want at the end of it. Turn the toggle off and you get a silent clip.