Skip to main content

FLUX 3 Image Prompting Guide

FLUX 3 Image.
Describe it like a photographer.

FLUX 3 Image is the image model from Black Forest Labs. It turns a written description into a photoreal image from a quick 768 square draft up to 4K, prints the words you put in quotes, edits and combines up to 10 reference images, and can research real places and products on the web before it draws. Every image on this page was made on PixelDojo with the prompt shown.

Ultra wide painted concept art of a giant humpback whale made of storm clouds swimming over a dark sea at sunset, with a white lighthouse and a tiny figure in a yellow coat on a cliff

Overview

On FLUX 3 Image the prompt carries everything. There is no negative prompt, no guidance slider and no seed, so the words you write are the whole control surface. That makes it simple to learn: describe the finished image the way you would brief a photographer, and add detail only where you want control.

A short prompt works, because the model expands it into a detailed description before it generates. Write more when you care about positions, light, materials or text. Black Forest Labs suggests a reliable order: the image type and medium first, then the subject, the setting, the light, and the framing and lens. Every prompt in this guide follows that order.

The same page does generation and editing. Leave the reference slot empty for text to image. Attach one to ten images and the prompt becomes an instruction: recolor this, replace that, put the product from image 2 into the room in image 1.

4K

Largest output

10

Reference images

15

Aspect ratios

Key Features

Photoreal Light and Materials
Concrete, oak, leather, wool

Photoreal Light and Materials

FLUX 3 Image renders surfaces the way light actually hits them. Name the light source and what it does, and the materials follow: low sun raking across polished concrete, a wool throw catching the edge of the window light.

Text You Put in Quotes
Headlines, dates, labels

Text You Put in Quotes

Black Forest Labs highlights accurate text rendering in multiple languages. Put the exact words in quotation marks with the case you want, and say where each block sits and how it looks. All three strings on this poster printed as written.

Edits That Keep the Rest
One reference, one instruction

Edits That Keep the Rest

Attach an image and write the change as an instruction. Name the target, say what it becomes, then list what stays. This edit recolored a product with a hex code and swapped the stone under it, and the carafe, steam and droplets did not move.

Up to 10 References
Composite products into scenes

Up to 10 References

Give each image a role and refer to it by position. Here image 1 is a cabin interior and image 2 is a product shot. The product landed on the table in the room's own late afternoon light, with the room untouched.

Web Grounding
Research before it draws

Web Grounding

With grounding on, the model looks your prompt up with web and image search before it generates. Use it for real products, real places and current events, like this blue hour view of the Ponte Vecchio in Florence.

15 Aspect Ratios, Up to 4K
21:9 through 9:21

15 Aspect Ratios, Up to 4K

Fifteen fixed ratios from ultra wide 21:9 to tall 9:21, and five sizes from 768sq to 4k. This 21:9 matte painting came back at 3136 x 1344 on the 2k setting for 1.5 credits.

Example Images

Each example shows the exact prompt that produced the result. Copy any prompt with one click.

Editorial Portrait

Editorial Portrait

1.5k · 4:5 · 1 credit

Editorial portrait photograph shot on medium format film. A ceramicist in her seventies with cropped silver hair and a rust colored linen work jacket sits on a wooden stool in her studio, clay dusted hands folded on one knee, looking straight into the lens with a calm half smile. Shelves of unglazed bowls fade into shadow behind her. Soft north window light from the left, deep falloff on the right, fine grain of warm color negative film, 80mm lens at f/2.8, real skin texture.

The first sentence sets the medium, the next ones give a specific person one pose, and the light gets a direction on both sides: "soft north window light from the left, deep falloff on the right". Naming a lens and film grain is what moves the result from clean digital to editorial.

Typographic Poster

Typographic Poster

1.5k · 2:3 · 1 credit

Full color scan of a screen printed concert poster in a mid century Swiss graphic style. A huge bold red sans serif headline across the top reads "NIGHT TIDES". Below it, a flat graphic illustration of a white heron standing in dark teal water under a giant orange moon. At the bottom, two centered lines of small black capitals read "LIVE AT THE BOATHOUSE" and "OCTOBER 18, DOORS 8 PM". Cream paper with fine grain and slight ink misregistration.

Each string is quoted exactly, with its size, color, case and position: headline across the top, two centered lines at the bottom. Describing the print process ("screen printed", "ink misregistration") gives the type a texture instead of a flat digital finish.

Overhead Food

Overhead Food

1.5k · 4:5 · 1 credit

Overhead food photograph of a black cast iron skillet of shakshuka with four glossy poached eggs in a bubbling spiced tomato sauce, crumbled feta, torn cilantro and a swirl of green olive oil. A torn sourdough loaf, a small bowl of flaky salt and a crumpled linen napkin sit beside it on a weathered pale blue wooden table. Soft daylight from a window at the top of the frame, shot straight down with a 50mm lens.

Counts and props are honored, so list them: four eggs, the loaf, the salt, the napkin. Placing the window "at the top of the frame" tells the model where the light comes from in a straight down shot, which is the part people usually leave out.

Night Market Candid

Night Market Candid

1.5k · 3:2 · 1 credit

Candid street photograph on 35mm film at night in a rainy open air market. A cook in a white apron tosses noodles in a flaming wok under a string of bare bulbs, the flames lighting his face and the rising steam. A blurred passerby with a clear umbrella crosses the foreground, wet pavement reflects red and amber lantern light. Grainy tungsten film look with halation around the lights, 35mm lens, shallow depth of field.

A candid needs one action beat and something in the way. The tossed noodles freeze the moment and the blurred passerby in the foreground sells the "caught, not posed" feel. "Halation around the lights" is a concrete film look that beats a vague word like "cinematic".

Product Hero (the reference)

Product Hero (the reference)

1.5k · 3:2 · 1 credit

Studio product photograph of a matte sage green ceramic pour over coffee dripper resting on a clear glass carafe on a slab of dark wet slate, a thin stream of coffee falling into the carafe, a wisp of steam curling up, water droplets beading on the slate. Single hard key light from the upper left with soft fill, a deep charcoal backdrop, clean negative space on the right, 100mm macro lens, crisp ceramic glaze texture.

Ask for the layout you need later: "clean negative space on the right" leaves room for a headline. This shot is the before image for the edit below and for the cabin composite in the strengths above.

Reference Edit (the after)

Reference Edit (the after)

1.5k · auto ratio · image 1 = the product hero above · 1 credit

Change the dripper and its saucer in image 1 to a glossy cobalt blue glaze (#1f4fd8) and change the dark slate slab to pale cream travertine stone. Keep the glass carafe, the coffee stream, the steam, the water droplets, the lighting and the framing exactly the same.

Compare it with the card above: only the glaze and the stone changed. The pattern is the one Black Forest Labs recommends for edits: name the target, say what it becomes (a hex code when the color matters), then list what stays. Leave the aspect ratio on auto so the output keeps the frame of image 1.

Prompting Tips

Set the aspect ratio yourself for text to image

Auto follows the first reference image. With no reference it gives you a square, so pick 16:9, 3:2, 4:5 or whatever fits the job, and make the prompt agree with the frame: an ultra wide frame needs something to fill the width.

Describe what you want, not what you do not

There is no negative prompt. Replace each "no" with what should be there instead: "an empty promenade" rather than "no crowds", "clean unmarked surfaces" rather than "no text", "plain studio backdrop in one color" rather than "no clutter".

Concrete beats labels

Words like "aesthetic" or "professional" give little direction. Name the medium and what it does to the picture: a flash snapshot with hard shadows, grainy film with halation, a flat screen print with ink misregistration.

Tie each hex code to one object

For an exact color, write the hex code next to the single thing it belongs to: "the dripper and its saucer ... (#1f4fd8)". A loose palette of hex codes leaves the model guessing where each one goes.

Quote text with the case you want

Put every word that must appear in quotation marks, exactly as it should read, and say where it goes. For several blocks, describe each one separately. Proofread small print before you publish.

Draft small, then go big

768sq, 1k and 1.5k all cost 1 credit, so shape the prompt there. Move to 2k or 4k for the final. A new resolution is a new generation, so the composition can shift; keep the prompt the same and judge the new frame on its own.

Switch grounding off for invented worlds

Grounding is on by default and is there for real places, products and current events. For pure fiction, like a whale made of clouds, it has nothing to look up, so you can turn it off.

In edits, name the one thing that changes

If several things match ("the car" in a street of cars), point at one by position, size or color: "the white car parked on the right". Then list what must stay, so the rest of the frame is protected.

The Five-Step Prompt Framework

1. Image type and medium. One opening sentence that sets the look: "Editorial portrait photograph shot on medium format film", "Full color scan of a screen printed concert poster", "Epic concept art matte painting, ultra wide cinematic shot". This sentence decides what kind of picture you get before anything else is read.

2. Subject. Who or what, how they look, and the one thing they are doing: "a ceramicist in her seventies with cropped silver hair and a rust colored linen work jacket ... clay dusted hands folded on one knee". Visible details that tell your subject apart do more than adjectives.

3. Setting. Where it happens and what surrounds the subject: "shelves of unglazed bowls fade into shadow behind her", "a weathered pale blue wooden table". If two subjects relate to each other, say how.

4. Light. The source, the direction, the quality and what it does: "soft north window light from the left, deep falloff on the right", "late afternoon sun casts long window shadows across a polished concrete floor". This is the single most useful line for photoreal work.

5. Framing and detail. Shot size, angle, lens, depth of field and texture: "shot straight down with a 50mm lens", "tilt shift lens with straight verticals", "grainy tungsten film look with halation". The order is a habit, not a rule: if the frame is the point of the picture, as in an ultra wide shot, open with it.

Editing and Combining Reference Images

Attach one to ten images and the prompt becomes an edit. There is no edit mode, mask or strength setting: the instruction says whether you are changing a detail, restyling, or combining images. Each input should be at least 256 x 256 pixels and no larger than 16 megapixels.

Refer to images by position. The first image you attach is "image 1", the next is "image 2", and so on. With the aspect ratio on auto, the output keeps the frame of image 1, so attach the image whose shape you want first. Set a ratio only when the edit should change the frame, for example turning a square product shot into a 16:9 banner, and say what fills the new space.

For a single image edit, use three moves: name the target precisely enough that only one thing matches, say what it changes into, then list what stays. Our product edit above follows it exactly: "Change the dripper and its saucer in image 1 to a glossy cobalt blue glaze (#1f4fd8) ... Keep the glass carafe, the coffee stream, the steam, the water droplets, the lighting and the framing exactly the same."

For several images, give each one a role: a subject to keep recognizable, a product whose shape and label must hold, a setting everything goes into, or a style to apply. Then say how the pieces fit, including whose light wins. Our cabin composite used: "Put the sage green ceramic dripper and glass carafe from image 2 on the walnut table in image 1, beside the stack of books, lit by the warm late afternoon window light of image 1 with a soft shadow on the table. Keep the room, the leather chair, the window view and the camera angle of image 1 exactly the same."

Removing something is an instruction, not a negation: "Remove the cat from the sofa" works as an edit prompt. Season, weather and time of day changes can be very short: "Change this to winter".

Web Grounding: Real Places, Products and Events

FLUX 3 Image on PixelDojo has a Web Grounding switch, on by default. With it on, the model researches your prompt with web and image search before it generates. That is what you want when the picture has to match something that exists: a landmark, a specific product, a current event, a place you have never seen.

Write the real thing by its name and the model does the looking up. Our Ponte Vecchio prompt named the bridge, the city and the river, then spent the rest of the words on what only we could decide: blue hour, the rowing scull, the camera at water level, the 35mm lens.

Grounding adds a research step to the job, so a grounded image can take longer than an ungrounded one. For invented scenes and pure style work there is nothing to look up, and switching grounding off keeps the run focused on your description.

Settings Reference

SettingValuesNotes
Promptstring (required)What to create, or the edit to apply when reference images are attached. Image type, subject, setting, light, framing.
Reference imagesup to 10 (optional)Turns the job into an edit or composite. Refer to them as image 1, image 2. Each at least 256 x 256 and at most 16 MP.
Resolution768sq · 1k · 1.5k · 2k · 4k1 credit for 768sq, 1k and 1.5k; 1.5 credits for 2k; 8 credits for 4k. Default 1k. One image per run.
Aspect ratioauto · 21:9 · 2:1 · 16:9 · 3:2 · 7:5 · 4:3 · 5:4 · 1:1 · 4:5 · 3:4 · 5:7 · 2:3 · 9:16 · 1:2 · 9:21Auto follows the first reference image, and is square when there is none.
Web groundingOn / OffOn by default. Researches the prompt with web and image search before generating.
Output formatjpg · png · webpDefault jpg.

FAQ

FLUX 3 Image is the image generation and editing model from Black Forest Labs, the team behind FLUX. It makes images from a text prompt and edits or combines up to 10 reference images, from a 768 square draft up to 4K.