WAN 3.0.
30 seconds, sound included, tested prompt by prompt.
A prompting guide for WAN 3.0 built from generations run on the live model, not from a spec sheet. Stage structure for long clips, reference-image roles and exclusions, first and last frames, adaptive framing, camera language, and the one common technique that does not work here.
WAN 3.0 is still in invited testing and is not yet available to generate on PixelDojo. The techniques below were verified on the live model, and the preview page has real clips from the same test.
Overview
WAN 3.0 generates video from a text prompt, from up to ten reference images, or from a first and last frame. Sound is generated with the picture and arrives in the same file. Clips run from 2 to 30 seconds in a single generation, and the model can pick both the frame shape and the length for you.
This guide is written from testing on the live model rather than from a spec sheet. Every clip on this page is one of those test generations, shown at the resolution it came out of the model. Where a technique works, there is a template for it. Where a technique that works on other models does not work here, that is called out instead of quietly left in.
A bicycle wheel is trued from start to finish in a single clip. Written as three stages with an explicit end state for each. Sound is generated with the picture.
- 2 to 30 seconds in one generation, or -1 to let the model choose the length.
- Native audio: dialogue, music, room tone, and foley in the same file.
- Adaptive aspect ratio: the model reads the prompt and picks the frame shape.
- Up to 10 reference images, or a first frame and a last frame. Not both.
- No negative prompt. WAN 3.0 removed it, so exclusions go in the prompt itself.
What Transfers From Other Models
Prompting habits built on Seedance and similar models mostly carry over, with one clear exception. Each row below was tested as a matched pair against the live model: same subject, same seed, same duration, with only the technique under test changed.
| Technique | Result on WAN 3.0 | What we saw |
|---|---|---|
| Stages with explicit end states | Works, and matters most here | A 20-second three-stage prompt hit every stated end state, including the small ones: bench clear after stage one, tool still in hand at the end of stage two, stand empty at the end. The same events written as one run-on sentence produced a muddier clip that kept re-handling the same object. |
| Timestamp ranges | Works | Three consecutive four-second ranges each played in order, and each honoured its "only X remains" end state exactly. |
| Naming references and excluding what not to use | Works, but the test did not prove it is necessary | With two reference images and explicit exclusion lines, neither excluded element appeared. The control with no role language did not leak them either, so with an obvious person-plus-place split WAN 3.0 already composites sensibly. Treat exclusions as cheap insurance, not a proven requirement. |
| Observable behaviour instead of emotion words | Works | Named cues (fixed gaze, held breath, still fingers) held a steady performance. The abstract version drifted into a much broader, near-weeping expression and ignored the stated setting. |
| Translating a camera term into a visible change | Works, and is worth the words | Writing out what a rack focus should look like produced a clean focus pull from foreground leaves to the subject. The bare term "rack focus" produced no focus change at all. |
| Subject, action, scene, style, camera, audio ordering | Works, buys control rather than quality | A bare one-line prompt still produced a competent multi-shot clip. The structured version was not prettier, it was obedient: it executed the one camera move that was asked for instead of inventing its own coverage. |
| Bracket syntax for music, effects, dialogue, subtitles | Does not transfer | Across two matched pairs, the bracket form produced on-screen text zero times out of two. Plain sentences produced it once out of two. Brackets are not a WAN 3.0 convention, and on-screen text is unreliable however you ask for it. |
The short version: keep the structure, drop the punctuation. Stages, end states, named references, exclusions, and observable detail all work. Bracket markers do not.
Two honest caveats on the audio and text row. First, on-screen text is inconsistent on WAN 3.0 in general: our best result was a clean, correctly spelled title card from a plain sentence, but the same phrasing on a different scene produced nothing. Plain language is the better bet, not a guarantee. Second, we confirmed every clip carried an audio track but did not verify by ear that spoken dialogue matched the requested words, so treat the dialogue half of that row as unconfirmed rather than proven.
The Core Prompt Formula
Subject + Action or Event + Scene and Environment + Visual Style + Camera Movement or Cut + Audio
WAN 3.0 fills gaps confidently. Leave the camera unspecified and it will choose its own coverage, often cutting between two or three setups inside five seconds. That is fine when you want a finished-looking clip and unhelpful when you need one specific shot. Name the parts you care about and let it invent the rest.
<Subject> performs <primary action or event> in <scene and environment>. The visuals feature <visual style>. Use <shot size, camera angle, camera movement, or cuts>. Audio includes <dialogue, ambience, sound effects, or music>.
A woman pours hot water in slow spirals over a paper filter of ground coffee, then lifts the kettle away, in a narrow kitchen at sunrise. The visuals feature low warm sunlight from a window on the left, steam catching the light, matte ceramic and brushed steel surfaces. Begin on a medium shot of her hands at the kettle, then slowly push in to a close-up of the coffee bed blooming. Audio includes the trickle of water, the low hiss of steam, and quiet morning room tone.
Competent, and entirely the model’s own film. It picked the kitchen, cut between four setups, and ended on the carafe.
One push-in, as asked. Light from the stated side, steam catching it, ending on the coffee bed. Same seed as the bare version.
Neither clip is bad. That is the point: the structured prompt did not buy quality, it bought obedience. Use the formula when the shot has to match something in your head, and a one-liner when you are happy to be surprised.
Long Clips: Stages and End States
Thirty seconds in one generation is the headline feature, and it is also where prompt structure pays off most. Split the clip into consecutive stages. Give each stage one primary change and say what should be visible when it ends. The end state is the part that does the work: it gives the model a checkpoint to hit rather than a mood to sustain.
[Generation Goal] Generate a <video type>. The central subject is <subject>, and the primary event is <story summary>. [Stage 1] Initial state: <initial state of characters, props, and scene>. Primary event: <one primary action or event>. End state: <positions, ownership, or visible scene state>. [Stage 2] Initial state: <continues from Stage 1's end state>. Primary event: <one primary action or event>. End state: <positions, ownership, or visible scene state>. [Stage 3] Initial state: <continues from Stage 2's end state>. Primary event: <one primary action or event>. End state: <final visible state>. Keep <identity, clothing, layout, and lighting> consistent throughout.
[Generation Goal] Generate an instructional video in a small bicycle repair shop. The central subject is a mechanic, and the primary event is truing a bent wheel. [Stage 1] Initial state: a bent wheel lies flat on the workbench, and the truing stand beside it is empty. Primary event: the mechanic lifts the wheel and seats it in the truing stand. End state: the wheel stands upright in the stand, and the workbench surface is empty. [Stage 2] Initial state: the wheel stands upright in the stand. Primary event: the mechanic spins the wheel and tightens spokes with a spoke wrench. End state: the wheel is still in the stand and spins without wobbling. The spoke wrench is in the mechanic's right hand. [Stage 3] Initial state: the trued wheel is still in the stand. Primary event: the mechanic lifts the wheel out and hangs it on a wall hook. End state: the wheel hangs on the wall hook, and the truing stand is empty. Keep the mechanic's identity and clothing, the shop layout, and the afternoon light from the shopfront window consistent throughout.
The wheel gets picked up and put down repeatedly, the stand is used ambiguously, and the ending is hard to read as "hung on a hook".
Watch the end states land: bench clear once the wheel is in the stand, wrench still in hand at the end of the spoke work, stand visibly empty once the wheel is on the wall.
- One primary change per stage. Two changes in one stage tend to collapse into whichever is easier to render.
- Write end states as things a viewer could point at: an empty surface, an object in a named hand, a door that is now closed.
- Carry the previous end state into the next initial state. It costs a line and keeps the geography stable.
- Put the continuity line last, naming identity, clothing, layout, and lighting.
Timestamps and Pacing
Use stages by default. Reach for timestamps when you need a specific beat to land at a specific moment, such as a handoff, an entrance, or a transition. Ranges should be consecutive and should not overlap.
0-4 seconds: an empty wooden display table in soft studio light. A hand places a white ceramic plate in the center. End state: the hand has left the frame and only the white plate remains. 4-8 seconds: the white plate is removed and a clear glass is placed in the center. End state: only the clear glass remains. 8-12 seconds: the clear glass is removed and a green ceramic vase is placed in the center. End state: only the green vase remains. Locked-off frontal camera throughout.
Plate, then glass, then vase. Each object is cleared and the hand leaves frame before the next range starts, exactly as the end states asked.
All three beats played in order and each one cleared the frame before the next object arrived. Treat a range as a time budget rather than an exact edit point, and do not ask for a rate of events such as three actions in one second.
Reference Images: Name Roles, State Exclusions
WAN 3.0 takes up to ten reference images. Because there is no negative prompt, everything you want kept out has to be said in the prompt. Bind each image to one job and add an exclusion whenever the image contains something you do not want, which is nearly always: a portrait carries a background, a location photo usually carries a person.
@Image 1 defines <Subject A>'s <face, hair, and clothing>. Do not use the image background. @Image 2 defines <Location>'s <structures, materials, and light>. Do not use the person in the image. <Subject A> performs <action> in <Location>. The visuals feature <visual style>. Use <camera treatment>.
@Image 1 defines the Knight's face, short dark red hair, and engraved silver armor. Do not use the image background. @Image 2 defines the Dock's wooden pilings, coiled rope, still water, and the mountains behind it. Do not use the person in the image. The Knight walks slowly along the Dock toward the camera and stops at the end of the boards. The visuals feature flat overcast daylight. Use a locked-off wide shot.
Knight from image 1, dock from image 2. No cherry blossoms from image 1’s background, no fisherman from image 2.
It also kept both exclusions out, and followed "walks toward camera" more literally. The honest read is that this test did not prove exclusions were needed here.
Image 1 was a portrait with a distinctive blossom background. Image 2 was a dock with a man leaning against a post. The version with roles and exclusions kept the knight and the dock and dropped both the blossoms and the man. So did the control, which is worth saying plainly: with two images that split cleanly into a person and a place, WAN 3.0 worked out the assignment on its own.
Keep writing the roles anyway. They cost a line, they remove the guesswork as soon as you go past two images or use references that overlap in subject matter, and they are the only way to express an exclusion at all now that the negative prompt is gone.
- Bind one image to one job. Do not write that images 1 through 4 define four characters, because that never says which is which.
- If several images show the same subject from different angles, say they are the same subject and state how many should appear in the output.
- Reference images and first or last frames cannot be combined. Sending both is rejected.
- Expect reference jobs to take much longer to render than text-only jobs. In testing a text-only clip returned in minutes and a two-image job took closer to half an hour.
First and Last Frames
Supply a first frame and the clip opens on it. Supply a last frame as well and the model generates the motion between the two. This is the mode to use when a shot has to start or land on something specific.
<Describe one continuous action that moves from the first frame to the last frame>. Keep <subject identity, clothing, prop positions, and camera framing> unchanged except where the action requires it. Do not <specific drift you want to prevent, for example: do not move the camera off the subject or leave the starting position>.
The opening frame reproduces the source image. By the end the subject has left the post and the camera has moved behind him, which the prompt never asked for.
Our test opened on the supplied frame precisely. It then went further than the prompt asked: the instruction was that the subject stays leaning against a post and turns his head, and by the end he had left the post and the camera had travelled behind him. A first frame anchors the opening, not the whole clip, so state plainly what must not change if you need the shot held.
Letting the Model Pick Shape and Length
Two options here have no equivalent on most other models. Set the aspect ratio to adaptive and WAN 3.0 reads the prompt to choose the frame shape. Set the duration to -1 and it chooses the length.
A vertical social ad for a stainless steel water bottle. The bottle stands on wet river stones, the camera rises slowly, and condensation runs down the side. Text-free, product-first framing for a phone screen.
The prompt says "vertical social ad" and nothing else about shape or length. The model returned 480 by 832, about six seconds.
That prompt never mentions dimensions or seconds. It returned a vertical clip of about six seconds. With a first frame supplied instead, adaptive matched the ratio of the supplied image. Say what the clip is for, in words, and the setting does the rest.
- Adaptive reads intent from phrases like vertical social ad, square feed loop, or widescreen trailer.
- Duration -1 is useful when the action has a natural length and you would rather not pad or clip it.
- Pin both settings explicitly when you are matching an existing edit, because inferred values will vary between runs.
Audio and On-Screen Text
This is the one place where habits from other models actively hurt. Bracket markers for music, sound effects, dialogue, and subtitles are not a WAN 3.0 convention. Across two matched pairs they produced on-screen text zero times out of two, while plain sentences produced it once out of two. The one success was a clean, correctly spelled title card that faded in and held.
Read that carefully, because it is a weaker claim than it first looks. Plain language is the better phrasing, but on-screen text is genuinely unreliable on WAN 3.0: the same plain wording that produced a title card on one scene produced nothing on another. Sound itself is dependable and arrives on every clip. If a caption or title is load-bearing for your edit, add it in post rather than hoping the model renders it.
A lighthouse keeper in a yellow raincoat pushes open a heavy door at the top of a lighthouse and steps out onto the gallery in wind and rain. Slow low strings play underneath. Wind gusts hard and rain rattles on metal. The keeper turns to the camera and says that the light is still turning. A subtitle reading Chapter One: The Keeper appears on screen.
A lighthouse keeper steps out onto the gallery in wind and rain.
(Slow, low strings play underneath)
<Wind gusts hard and rain rattles on metal>
The keeper says: {The light is still turning.}
【Chapter One: The Keeper】The scene is fine. No title card appears at any point, and the same happened on a second scene.
The best text result we got: "Chapter One: The Keeper" fades in, correctly spelled, and holds. The same phrasing on a cafe scene produced nothing, so this is the good case rather than the typical one.
- Describe music by mood and instrument, in a sentence.
- Describe sound effects as events that happen in the scene.
- Introduce dialogue with a plain speech tag naming the speaker.
- Ask for on-screen text by describing it: a subtitle reading X appears on screen. Put the exact wording in the sentence, then check the result and be ready to add it in post.
- Audio is generated either way. Turning it off does not reduce the cost of a generation.
Directing Performance
Emotion words set a direction and leave the rest open. WAN 3.0 fills that opening generously, and usually broader than you intended. Naming two to four observable cues holds the performance where you want it.
A young violinist sits alone backstage before a performance. Her fingers press hard against the neck of the violin and stop moving, her gaze fixes on the floor without blinking, her shoulders stay lifted, and she takes one short breath in and holds it. Medium close-up, static camera.
Starts neutral and escalates to a near-weeping expression, and quietly moves her from backstage onto a lit stage.
Fixed gaze, held breath, still fingers, lifted shoulders. One contained state, held for the whole clip, in the setting that was asked for.
The paired control said only that she looks nervous and tense. It drifted from neutral into a near-weeping expression by the end and relocated her from backstage to a lit stage. The cue-based version held one contained state for the whole clip and stayed where it was put.
The overall emotion shifts from <starting emotion> to <ending emotion>. After <triggering event>, <subject> first shows <immediate observable reaction>. Then, <eyes, brows, mouth, breathing, gaze, or hand movement> gradually <changes>. Finally, <subject> expresses <target emotion> through <restrained or explicit outward behavior>.
Camera Language
Plain shot sizes and camera moves work written as they are. Named techniques are less reliable on their own. The strongest single result in our testing came from writing out what a technique should look like rather than naming it.
| Type | Terms that work as written |
|---|---|
| Shot size | extreme wide shot, wide shot, medium shot, close-up, extreme close-up |
| Camera movement | push in, pull out, pan, lateral move, follow shot, orbit, tilt up, handheld shake, locked-off |
| Position and viewpoint | low angle, overhead view, first-person view |
Camera Term + Target Subject + Visible Change + Foreground and Background Relationship + Direction or Speed
Works: In a greenhouse, a gardener stands among hanging ferns. Rack focus: shift focus smoothly from the fern leaves in the foreground to the gardener in the background. The leaves gradually blur while the gardener's face changes from soft to sharp. The camera does not move. Does not: Rack focus in a greenhouse. A gardener stands among hanging ferns.
The gardener stays soft for the whole clip. No focus pull ever happens.
Leaves sharp and face soft at the start, face sharp and leaves soft by the end, camera locked. Same seed as the bare version.
The written-out version produced a clean pull: foreground leaves sharp and subject soft at the start, subject sharp and leaves soft by the end, with the camera locked. The bare term left the subject soft for the entire clip and never pulled focus. This was the largest single difference in the whole battery, and it costs one extra sentence.
Settings Reference
| Setting | Values | Notes |
|---|---|---|
| Duration | 2 to 30 seconds, or -1 | One generation, no stitching. -1 lets the model choose. A test at -1 returned about six seconds for a short product shot. |
| Resolution | 480P, 720P, 1080P | Uppercase. Test clips at 480P came back 832 by 480 at 30fps. |
| Aspect ratio | 16:9, 4:3, 1:1, 3:4, 9:16, adaptive | Adaptive infers the shape from the prompt, or matches the supplied frame when one is given. |
| Reference images | Up to 10 | Cannot be combined with a first or last frame. Renders considerably slower than text-only. |
| First and last frame | One each | Cannot be combined with reference images. The first frame anchors the opening only. |
| Audio | On or off | Native, in the same file. Cost is the same either way. |
| Negative prompt | Not supported | Removed in WAN 3.0. Put exclusions in the prompt. |
| Seed | Integer | Fix it when comparing two prompts, so only the wording changes. |
What WAN 3.0 Does Not Do
Several techniques worth knowing from other models have no equivalent here. Reaching for them wastes a generation.
- No video editing. There is no mode that takes a finished clip and changes one object, region, or sound inside it.
- No video extension. You cannot continue an existing clip forward or backward from its boundary frame. Use the 30-second ceiling instead of stitching.
- No multi-keyframe sequences. Control points are limited to a first frame and a last frame, not an ordered set.
- No negative prompt. Every exclusion has to be a sentence in the prompt.
- Reference images and first or last frames are mutually exclusive, so a shot cannot be both character-matched and frame-anchored in one generation.
Storyboard grids and blockout references were not tested. They are plausible through the reference-image path, but we have no evidence either way and would rather say so than guess.
Before You Generate
- Does the prompt state the subject and the primary action?
- If the clip runs longer than about ten seconds, is it split into stages with one change each?
- Does every stage say what should be visible when it ends?
- Does every reference image have one job and, where needed, an exclusion?
- Are exclusions written as sentences, given that there is no negative prompt?
- Is audio and on-screen text written in plain language rather than bracket markers?
- Is any on-screen wording spelled out exactly as it should appear?
- Are emotions expressed as observable behaviour rather than mood words?
- Is any named camera technique also described as a visible change?
- If the shot must hold, does the prompt say what must not change?
- Are the aspect ratio and duration pinned, or deliberately left adaptive and -1?