Skip to main content
Qwen 3 Pro Prompting Guide

Qwen 3 Pro.
Generate and edit in one model.

Qwen 3 Pro is Alibaba’s newest image model, and it works two ways from the same page. Write a prompt and it generates. Attach one to three reference images and the same prompt becomes an edit instruction. Sharp detail, close instruction following, and text that actually renders. 1 credit per image at 1K, 2 credits at 2K.

Qwen 3 Pro hero image

Overview

Most image tools make you pick a lane. One model generates, a different one edits, and moving between them means re-uploading, re-prompting, and losing the look you just found. Qwen 3 Pro collapses that into a single tool. Leave the image slot empty and your prompt is a generation. Fill it with one to three references and the same prompt is read as an edit instruction against them.

It is built for the cases where precision matters: detail that holds up at full size, instructions that get followed literally rather than loosely interpreted, and words in the image that come back spelled correctly. Pick a 1K size for 1 credit per image, a 2K size for 2, or leave the size on Auto and the model picks for you (Auto matches the input image when you are editing). Up to 6 images per run.

1 to 2 cr

Per image

16

Size presets

3

Reference images

Key Features

Generate and Edit in One Tool
One prompt box, two modes

Generate and Edit in One Tool

No image attached means text to image. One to three images attached means editing, and your prompt becomes the instruction. Nothing to switch, no second tool to open. Generate a scene, attach the result, and refine it without leaving the page.

Close Instruction Following
Says what it does

Close Instruction Following

Qwen 3 Pro reads a prompt as a specification rather than a mood board. Counts, placements, and relationships (three bottles, sign on the left, hat removed) land far more often than they do on models that treat the prompt as loose inspiration. Long prompts do not get quietly averaged away.

Text That Renders
Signage, packaging, posters

Text That Renders

Put the exact string in double quotes and it comes back spelled correctly, in the material you asked for. Painted window signs, chalkboards, product labels, and book covers all hold up. This is one of the model’s strongest dimensions, so lean on it.

Multi-Image Fusion
Up to 3 references

Multi-Image Fusion

Attach up to three images and refer to them by position: the character from image 1, the environment from image 2, the lighting from image 3. Useful for putting a consistent subject into a new scene, or borrowing a palette without describing it.

Two Prices, No Surprises
1 credit or 2

Two Prices, No Surprises

Every 1K size costs 1 credit per image. Every 2K size costs 2. Auto bills at the 2K rate because the model can pick anything up to the cap. That is the whole pricing table, so drafting on 1K and finishing on 2K is an easy habit to keep.

Example Images

Each example shows the exact prompt that produced the result. Copy any prompt with one click.

Photoreal Portrait

Photoreal Portrait

3:4 at 2K (1080*1440) · 2 credits

Photoreal portrait of a ceramicist in her studio, dried clay dust on her forearms, hands resting on a half-thrown bowl, soft window light from the left, shelves of unglazed pots blurred behind her, warm neutral palette, fine skin texture, 85mm look

Name the light before you name the camera. "Soft window light from the left" does more work here than any lens number, and the occupational detail (clay dust, half-thrown bowl) gives the model a reason for the pose instead of a stock expression.

Cinematic Landscape

Cinematic Landscape

16:9 at 2K (1920*1080) · 2 credits

Wide cinematic shot of a black sand beach at dawn, basalt sea stacks standing in low mist, a single figure in a red jacket walking the shoreline for scale, cold blue light with a thin warm band on the horizon, long exposure water, high detail

One small figure for scale is the oldest landscape trick there is, and Qwen 3 Pro places it thoughtfully rather than filling the frame with it. Pair a structural element (basalt stacks) with a two-part color plan (cold blue, warm horizon band) and the scene reads as one image rather than a list.

Signage and In-Image Text

Signage and In-Image Text

3:2 at 2K (1536*1024) · 2 credits

Storefront of a small coffee roaster at golden hour, hand painted window sign reading "SLOW POUR" in worn gold leaf, a chalkboard beside the door reading "Single Origin Today", weathered brick facade, bicycle leaning by the entrance, warm morning light

Quote the exact string and say what it is painted on. Two separate text elements (window sign plus chalkboard) is comfortably within range. Describing the material of the lettering, like worn gold leaf, is what keeps it from looking like a font pasted over a photo.

Product Shot

Product Shot

1:1 at 2K (1536*1536) · 2 credits

Studio product photograph of a matte black stainless steel water bottle standing on a pale concrete plinth, one soft key light from the upper right, subtle gray gradient background, crisp specular edge along the left side, generous negative space above, advertising quality

Product work is a lighting brief, not a subject description. One named key light, one named background, and one named highlight is usually enough. "Generous negative space above" leaves you room to drop a headline in later.

Stylized Illustration

Stylized Illustration

2:3 at 1K (768*1152) · 1 credit

Flat vector illustration of a night market street, two color risograph palette in teal and coral, bold simple shapes with no gradients, paper grain texture, paper lanterns strung overhead, small crowd rendered as silhouettes, generous margins

For stylized work, say what to leave out. "No gradients" and "silhouettes" are constraints, and this model honors constraints well. Illustration iterates cheaply at 1K, so explore the palette there and only spend 2 credits on the version you like.

Multi-Subject Scene

Multi-Subject Scene

16:9 at 2K (1920*1080) · 2 credits

Three chefs plating dishes along a stainless steel pass in a busy restaurant kitchen, steam rising into overhead heat lamps, tickets clipped along the rail, each face distinct and in focus, candid documentary photography, warm practical lighting

Give the count and give the arrangement. "Three chefs along a pass" is a layout instruction, and the model follows counts closely. Adding "each face distinct and in focus" is the guard against the crowd blur that multi-subject prompts usually drift toward.

Prompting Tips

Quote the exact text

Put every string you want rendered inside double quotes, and say what it is written on: painted on glass, chalk on slate, embossed on a label. Unquoted text gets treated as description and often comes back paraphrased.

Draft at 1K, finish at 2K

All eight 1K sizes cost 1 credit per image. Explore composition and wording there, then rerun your winner at the matching 2K size for the final. Same aspect ratios in both tiers, so nothing about the framing changes.

Use negative prompts sparingly

A short list of genuine problems works. A long wall of negatives strips detail you wanted and makes results harder to reason about. If something keeps appearing, first try describing what should be there instead.

Set a seed to compare fairly

Fix the seed when you want to test one prompt change at a time. Same seed with the same inputs gives a closely related result, not a byte identical one, which is enough to see what your edit did. Leave it empty for variety.

Ask for up to 6 at once

The output count goes to 6 and credits are charged per image. Six at 1K costs 6 credits and gives you a real spread to choose from, which beats running the same prompt six times and losing track of which was which.

Reference order matters

Images are numbered in the order you attach them, so "image 2" means the second one you added. When you leave the size on Auto, the output takes its aspect ratio from the LAST image in the list. Put the shape you want the output to be at the end.

A Framework for Text to Image Prompts

Prompts for this model work best built in four passes. Write the first draft in one line, then add each layer only if it changes the picture. Padding a prompt with adjectives that mean nothing to the scene tends to dilute the parts that do.

1. Subject. Say who or what, and one concrete thing about them. "A ceramicist" is a category. "A ceramicist with clay dust on her forearms, hands on a half-thrown bowl" is a picture. One specific physical detail is worth five adjectives.

2. Scene and setting. Where this is happening and what is around it. Name two or three supporting objects rather than describing a whole room. Background elements are how the model decides depth, so "shelves of unglazed pots blurred behind her" also tells it where the focal plane sits.

3. Style and lighting. Pick the register (photoreal, flat vector, editorial illustration) and then name the light: direction, quality, and time of day. "Soft window light from the left at morning" beats a lens number nearly every time. Add the camera language only when it changes something.

4. Text in the image. Put any words you want rendered in double quotes exactly as they should appear, then say what surface they are on and what condition they are in. "Hand painted window sign reading “SLOW POUR” in worn gold leaf" gives the model the string, the medium, and the wear all at once.

Editing With Reference Images

Attach one to three images and the prompt stops being a description and starts being an instruction. Write it as a command, not a caption: "make his shirt red", "remove the hat", "replace the background with a snowy street at night". Describing the finished picture instead of the change is the most common reason an edit comes back looking like a fresh generation.

With more than one image, refer to them by position: "put the character from image 1 into the environment from image 2", or "apply the color palette of image 2 to image 1". Numbering follows attachment order, so rearranging your uploads rewrites your prompt. Accepted formats are JPG, PNG, BMP, TIFF, WEBP, and GIF, up to 10 MB each.

Leave the size on Auto for edits. Auto matches the input image, so the result comes back at the same shape you started with and nothing gets cropped or letterboxed. Picking an explicit size is for when you deliberately want a different crop, for example turning a square product photo into a 16:9 banner.

Keep prompt enhancement off when the edit has to be exact. It is off by default for that reason: enhancement rewrites your instruction with a language model first, which is useful for a three word idea and actively harmful when you have specified precisely which button on the jacket to change. Turn it on for exploration, off for production.

Settings Reference

SettingValuesNotes
PromptString, requiredWhat to create, or the edit instruction when reference images are attached.
Reference images0 to 3None means text to image. One to three means editing. JPG, PNG, BMP, TIFF, WEBP, GIF, up to 10 MB each.
Size8 sizes at 1K, 8 at 2K, or AutoBoth tiers cover 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, and 21:9. Auto lets the model pick, and matches the input image when editing.
Outputs1 to 6 per runCredits are charged per image, so 4 images at 2K costs 8 credits.
Negative promptString, optionalWhat to keep out. Short and specific beats long and generic.
Prompt enhancementBoolean, off by defaultRewrites your prompt before generating. Leave it off when an edit has to be followed exactly.
WatermarkBoolean, off by defaultAdds a small mark in the bottom right corner.
Seed0 to 2147483647, optionalSame seed with the same inputs gives a closely related result. Leave empty for a random one.
Pricing1 credit at 1K, 2 credits at 2KAuto bills at the 2K rate because the model can pick any size up to the cap.

FAQ

Any of the eight 1K sizes costs 1 credit per image. Any of the eight 2K sizes costs 2. Leaving the size on Auto costs 2 credits per image. Credits are charged per image, so a run of 6 at 1K is 6 credits and a run of 6 at 2K is 12.