WAN 3.0 makes 30 second videos.
With sound.
We have early access while it is still in testing. Every clip below is one generation, straight out of the model, sound included.
Written as a screenplay, generated as one clip
Two characters, two exchanges of dialogue, a rescue at a gap in the stonework, and a payoff on the ledge. Thirty seconds, no cuts, sound generated with it.
New in 3.0
30 seconds in one generation
Any length from 2 to 30 seconds, generated as a single clip. A scene with a beginning, middle, and end no longer has to be stitched together from fragments.
Sound in the same file
Dialogue, music, room tone, and foley are generated with the picture and arrive in the same mp4. Turning sound off costs exactly what leaving it on costs.
It picks the frame shape
Set the aspect ratio to adaptive and the model reads the prompt. We asked for a vertical ad and got 720x1280. We asked for a square product loop and got 960x960. Neither request mentioned dimensions.
Ten images, five clips, five sounds
That is what fits in one request as reference material. Pull a face from one image, a location from another, a voice from an audio file, and describe how they meet.
It can pick the length too
Ask for a duration of -1 and the model decides. We described a paper boat washing down a gutter and it returned 20 seconds, a length we never asked for.
Start and end frames
Hand it a first frame and a last frame and it generates the motion between them, which is the mode to reach for when you need a shot to land on something specific.
The clips
Each one was written to test a single claim on this page. Sound on.
Two chefs, mid-service
Two characters, four lines of dialogue, one generation.
Spoken dialogueThe letter
A letter goes from postbox to doorstep without a cut.
30-second single takeSunrise run
The prompt said vertical social ad and named no dimensions. It came back 720x1280.
Adaptive aspect ratioPerfume loop
The prompt said square feed loop and named no dimensions. It came back 960x960.
Adaptive aspect ratioBasement show
The music was generated with the picture, not laid over it.
Native audioRain on canvas
Rain on canvas, thunder, and a fire outside the frame, all from the prompt.
Native audioDashboard over the shoulder
An interface animating on screen, generated at 1080P.
Screen and text renderingHoney pour
Honey over a warm biscuit, macro, 1080P.
Photorealistic detailReference to video
One still of a person in, a moving shot of that person out.
Reference to videoPerson plus place
Two references in one request: the person from one, the forest from the other.
Multi-referenceFirst frame to last frame
Two stills set the start and the end. The model generated the middle.
First and last framePaper boat
Duration set to -1. The model chose 20 seconds.
Intelligent durationOur own preset prompts
Twelve prompts from our preset library, written for other video models, run through WAN 3.0 at 15 seconds with nothing reworded.
Character walk
Stylized character walking toward camera in a neon-lit alley.
Preset libraryDance loop
Vertical dance performance against a colorful studio backdrop.
Preset libraryPet action
Slow-motion frisbee catch in a sunny park.
Preset libraryProduct spin
Square 360-degree product rotation on a seamless backdrop.
Preset libraryReveal shot
Slow tilt-up hero reveal at golden hour.
Preset libraryCinematic drone
Aerial pass over coastal cliffs at golden hour.
Preset libraryFashion runway
Tracking shot down a runway with slow-motion fabric movement.
Preset libraryKitchen cooking
Top-down knife work on a worn wooden board.
Preset libraryLifestyle vlog
Handheld morning-routine B-roll in warm natural light.
Preset libraryProduct explosion
Slow-motion disassembly and reassembly in black space.
Preset libraryProduct rotation
Premium product turning under a studio key light.
Preset libraryTravel timelapse
A European plaza from dawn to night, light trails and moving crowds.
Preset libraryWhen you can use it
WAN 3.0 is still in testing and is not in the tool list yet. Everything else in our video lineup is live now.