Skip to main content
Seedance 2.0 Prompting Guide

The Next Generation
of AI Video

Cinematic output with native audio, real-world physics, and director-level camera control. Accepts text, image, audio, and video inputs — up to 12 assets in a single generation.

Overview

Seedance 2.0 is the most advanced Seedance video generation model. It uses a unified multimodal architecture that processes text, images, video clips, and audio together — generating cinematic video with native audio in a single pass.

What sets it apart is the combination of multi-shot storytelling, precise camera control, realistic physics simulation, and joint audio generation. One prompt can produce a multi-camera sequence with synced sound effects, dialogue, and music.

15s

Max video length

12

Input assets per generation

A+V

Native audio-video output

Quality Tiers

Seedance 2.0 comes in three quality tiers on PixelDojo. Pick the tier in the quality selector on the tool page — same prompts, same workflow, different speed-cost-fidelity trade-off.

Mini

The budget tier — about half the credits of Fast (1 credit/sec at 480p, 2 at 720p). Ideal for high-volume work: iterating on prompts, batch marketing assets, UGC-style content, and quick drafts before a final high-quality render.

Fast

The default balance of speed and quality (2 credits/sec at 480p, 4 at 720p). Great for everyday generations where turnaround matters.

High

Maximum fidelity, and the only tier with 1080p and 4K output. Use it for final renders, client deliverables, and anything headed to a big screen.

A workflow that stretches credits: draft and iterate on Mini at 480p until the motion and framing are right, then re-run the winning prompt on High for the final cut.

Key Features

Turn on audio to hear the native sound generation. Every example below was generated in a single pass with no post-production.

Advanced Cinematography

Director-Level Camera Control

Dolly zooms, rack focuses, tracking shots, POV switches, and smooth handheld movement. Describe the shot you want, and the camera executes it.

Real-World Physics

Action That Feels Real

Fight scenes, vehicle chases, explosions, falling debris. Collisions have weight, fabric tears realistically, and characters move with physical believability even in high-action sequences.

Audio-Video Joint Generation

Cinema-Grade Sound, Built In

Seedance 2.0 generates audio natively alongside video. Music carries deep bass and cinematic warmth. Dialogue is clear with precise lip-sync. Sound effects land exactly on cue.

Example Videos

Each of these videos was generated from a single text prompt. Copy any prompt to use as a starting point for your own generations.

High-Action Chase with Dynamic Tracking

Multi-camera action with crowd physics

Camera follows a man in black sprinting through a crowded street, a group chasing close behind. The shot cuts to a side tracking angle as he panics and crashes into a roadside fruit stall, scrambles to his feet, and keeps running. Sounds of a frantic crowd

Multi-shot action sequences work best when you describe camera angle changes alongside physical interactions and ambient audio cues.

Martial Arts Choreography in Nature

Complex multi-character combat with environmental interaction

A spear-wielding warrior clashes with a dual-blade fighter in a maple leaf forest. Autumn leaves scatter on each impact. Wide shot pulls into tight close-ups of parrying blades, then cuts to a slow-motion overhead as both leap into the air

For choreographed action, layer environmental reactions (scattering leaves, dust) with camera movement transitions and slow-motion beats.

Long-Take Spy Thriller

Continuous camera with character tracking and reveals

Spy thriller style. Front-tracking shot of a female agent in a red trench coat walking forward through a busy street, pedestrians constantly crossing in front of her. She rounds a corner and disappears. A masked girl lurks at the corner, glaring after her. Camera pans forward as the agent walks into a mansion and vanishes. Single continuous take, no cuts

Specify "single continuous take" and describe spatial transitions (rounding corners, entering buildings) to get long unbroken shots.

Multi-Shot Creative Commercial

Multi-cut commercial with text overlay and varied angles

15s commercial. Shot 1: side angle, a donkey rides a motorcycle bursting through a barn fence, chickens scatter. Shot 2: close-up of spinning tires on sand, then aerial shot of the donkey doing donuts, dust clouds rising. Shot 3: snow mountain backdrop, the donkey launches off a hillside, text 'Inspire Creativity, Enrich Life' revealed behind it as dust settles

Structure multi-shot commercials with explicit shot numbers, timings, and camera angles. Include text overlay instructions directly in the prompt.

Overdescription with Quality Anchors

Dense detail and fidelity keywords for maximum realism

A weathered fisherman mending nets on a sun-bleached wooden dock at golden hour, amber sunlight catching salt spray in the air, gnarled hands threading frayed twine through sun-bleached mesh, seagulls wheeling overhead against a lavender sky streaked with coral clouds, a distant lighthouse beam sweeping across choppy slate-blue waters. Hyper-realistic, 8k. Sounds of creaking dock wood, gentle waves lapping against barnacle-crusted pilings, and distant seagull calls.

Pack every prompt with sensory detail — textures, colors, light quality, sounds. Add "hyper-realistic, 8k" as quality anchors to push the model toward maximum fidelity.

Cinematographer Reference Style

Naming a DP to guide lighting, framing, and mood

Roger Deakins-style cinematography. A lone figure walks across a vast cracked salt flat at sunrise, their silhouette casting a razor-thin shadow across the white earth. Camera mounted low, slow dolly forward tracking the figure's boots. Warm amber backlight flares into the lens. The figure stops and looks toward the horizon. Hyper-realistic, 8k. Sounds of wind howling across the empty landscape and distant rumbling thunder.

Reference a cinematographer or director by name to steer the visual style. The model adapts lighting, lens choices, and camera movement to match their signature look.

Time-Coded Multi-Shot Sequence

Bracket timestamps for precise scene transitions

[0-5s]: Wide establishing shot of a neon-lit Tokyo alley at night, rain falling steadily, camera slowly dollying forward past glowing signs and steam vents. [5-10s]: Medium shot inside a tiny ramen shop, a chef ladles rich broth into a bowl with practiced precision, steam rising into warm overhead light. [10-15s]: Extreme close-up of the finished ramen bowl placed on the counter, chopsticks snapping apart, sounds of sizzling pork, bubbling broth, and the quiet murmur of the shop.

Use [0-Xs]: bracket notation to give the model explicit temporal structure. Each segment gets its own camera angle, subject, and audio cues for precise multi-shot control.

Wide-to-Close-Up Progression

Natural cinematic shot progression for momentum

A street violinist performing in a cobblestone European square at golden hour. Wide shot establishing the scene with passersby and autumn chestnut trees. Camera pushes in to a medium shot framing the musician's upper body and violin. Swift dolly zoom into an extreme close-up of fingers dancing across the strings, rosin dust catching the last rays of afternoon sunlight. Sounds of a melancholic violin melody echoing off sandstone walls.

Structure shots from wide establishing to medium to extreme close-up. This natural cinematic progression creates momentum, drawing the viewer deeper into the scene.

Prompting Tips

Structure multi-shot sequences

Label each shot with a number and timestamp. Include camera angle, subject action, and cut type: "Shot 1 (0-3s): Wide angle..."

Use explicit camera language

Call out movements: dolly, pan, tracking, crane, whip pan, slow push-in. Seedance responds to film terminology.

Add audio cues in the prompt

Describe the soundscape: "sounds of a frantic crowd", "autumn leaves rustling", "engine roar". Audio is generated natively.

Specify pacing and timing

Include temporal cues: "slow-motion overhead", "15s commercial", "single continuous take, no cuts" for control over rhythm.

Layer environmental reactions

Describe how the environment responds to action: scattering leaves, rising dust, shattering glass. It sells physical realism.

Mix input types for control

Combine images for visual style, audio clips for soundtrack, and text for scene direction. The model reads each input's role automatically.

Overdescribe the scene

Pack detail into every prompt. Instead of "a car chase," describe the rain-slicked streets, neon reflections, headlight beams cutting through mist. Seedance rewards specificity.

Use quality anchors

Add phrases like "hyper-realistic, 8k" as fidelity indicators. These push the model toward maximum visual quality and sharpness.

Try time-coded brackets

For multi-shot sequences, use [0-5s]: ... [5-10s]: ... bracket notation. The model recognizes this as explicit temporal structure.

Reference cinematographers

Name a director of photography like Deakins, Lubezki, or Kurosawa for style guidance. The model adapts lighting, framing, and movement to match.

Progress from wide to close-up

Structure shots from wide establishing to medium to extreme close-up. This natural cinematic progression creates momentum and draws the viewer in.

Label reference inputs

When using multiple images, videos, or audio clips, name them in the prompt as "Image 1", "Video 1", "Audio 1". See the Reference Assets section below for the full labelling rules.

Action Sequence Template

Use for multi-camera chase, fight, or high-energy scenes.

Camera follows [subject] through [environment], [action description]. The shot cuts to a [angle] as [secondary action]. [Environmental reaction]. [Audio cue].

Commercial / Ad Template

Use for multi-shot product or brand videos.

[Duration] commercial. Shot 1: [angle], [subject action]. Shot 2: [close-up detail], then [wide shot]. Shot 3: [final reveal], text '[tagline]' revealed as [transition].

Long-Take Narrative Template

Use for continuous single-take storytelling.

[Genre] style. [Camera movement] of [character description] moving through [environment]. [Character action]. [Secondary character reveal]. Single continuous take, no cuts.

Time-Coded Sequence Template

Use for precisely timed multi-shot transitions with bracket notation.

[0-Xs]: [Wide/medium/close], [camera movement], [subject action], [lighting]. [Xs-Ys]: [Cut type], [new angle], [new action]. [Ys-Zs]: [Final shot], [resolution moment], [audio cue].

Reference Assets

Every asset you upload gets a label the model can read. Write them in title case with a space before the number: Image 1, Video 1, Audio 1. No @ prefix, no brackets, no underscores. The numbers follow upload order.

Give each asset one job and one named subject. Then say which attributes carry over. An asset with two jobs pulls the model in two directions and usually loses both.

One role per asset

Each upload does one thing: a character, a location, a prop, a style, a music bed. Do not ask Image 1 to supply both a face and a background.

Name the subject

Attach a name to the role, not just a description. "Image 1: Courier" gives you a word you can reuse in every shot line.

List what is inherited

Say which attributes travel: face, hairstyle, garment, colour, texture, framing. "Inherit the face and the olive jacket" is a much tighter instruction than "use this person".

Short binding for simple scenes

For a one subject shot you can bind inline: "Courier@Image 1". For anything with several assets or several shots, define the role once up top and use the role name after that.

Reuse the role name

Once the Courier is defined, every shot line says Courier. Switching back and forth between the name and the asset label invites identity drift.

Keep the asset count honest

Only upload what the prompt actually addresses. An unreferenced asset still influences the result, and usually not in the direction you wanted.

Asset Role Block

Open every multi-asset prompt with a block like this.

Image 1: Courier. Inherit the face, the braided hair, and the olive jacket.
Image 2: Location. Inherit the rain-soaked alley and its neon signage.
Audio 1: Music bed. Inherit tempo and mood only.

Prompt Structure

Long prompts read better when they follow the order a production runs in. Roles first, then the goal, then the shots, then the direction that applies to everything, then the rules.

  1. Asset roles. What each upload is and what it hands over.
  2. Generation goal. One sentence on what you want built and what must stay consistent.
  3. Shot sequence. Shot 1, Shot 2, Shot 3, in the order they play.
  4. Global direction. Style, lighting, colour, and the sound bed that runs under the whole clip.
  5. Constraints. What must never appear.

Shot-level fields

Inside each shot, use these five fields. Drop any field you have nothing to say about. An empty label is noise, and the model reads it as one.

Camera

Shot size plus one primary movement. Wide with a slow dolly in. Close-up, static. Save a second movement for the next cut.

Subject action

What the named subject physically does in this shot, in the order it happens.

Space

Where the subject sits in frame, or how that position changes. Enters from frame right. Crosses to the far wall.

Audio

Dialogue, effects, and music for this shot only. Anything that runs the whole clip belongs in global direction.

End state

The visible state the shot finishes on. The next shot inherits it, so this is what keeps a cut from feeling like a jump.

One move per shot

Stacking a pan, a push, and a crane into one shot muddies all three. If you need a second move, cut and start a new shot.

Dialogue and text markers

Four bracket styles separate the kinds of sound. Using the right one keeps spoken lines out of the sound effects and keeps stage directions out of the dialogue.

  • { curly braces } hold words a character actually speaks.
  • (full-width parentheses) hold music direction such as genre, tempo, and mood.
  • <angle brackets> hold foley and ambience such as footsteps, rain, or room tone.
  • 【black lenticular brackets】 hold text you want rendered on screen. Leave these out entirely when you want a clean frame.

Three-Shot Reference Prompt

A full prompt in production order, with asset roles, shot fields, and markers.

Asset roles:
Image 1: Courier. Inherit the face, the braided hair, and the olive jacket.
Image 2: Location. Inherit the rain-soaked alley and its neon signage.
Audio 1: Music bed. Inherit tempo and mood only.

Goal: Reference the character identity from Image 1 and the setting from Image 2 to generate a three-shot night delivery scene. The Courier stays the same person in every shot.

Shot 1
Camera: Wide shot, slow dolly in.
Subject action: The Courier swings a leg off the bike and hunches both shoulders against the rain.
Space: Enters from frame right, stops under the neon sign.
Audio: <rain drumming on a metal awning>, (low pulsing synth, slow tempo)
End state: The Courier stands still, package pressed flat against the chest.

Shot 2
Camera: Medium shot, slight push in.
Subject action: The Courier raises the package chest high, then knocks twice with the free hand. Water runs off the knuckles.
Space: Turns to face a steel door on the left wall.
Audio: <two firm knocks, a deadbolt turning>
End state: The door opens a hand width. Warm light falls across the Courier's face.

Shot 3
Camera: Close-up, static.
Subject action: The Courier tips the chin up and lets out a short breath, then half smiles.
Audio: {Last one tonight.}
End state: The Courier holds the smile as the light widens.

Global direction: Neo-noir night exterior. Wet asphalt, magenta and cyan bounce light, shallow depth of field. Music stays under the dialogue throughout.

Constraints: One person on screen at all times. No duplicated Courier. No style change between shots. No subtitles, no on-screen text, no logos, no watermark.

Task Types

The opening verb tells the model what kind of job this is. Generation, editing, and extension want different phrasing, and mixing them is the fastest way to get a result that ignores half the prompt.

Reference generation

Open with "Reference [dimension] from Image 1 to generate...". Name the dimension you are borrowing: identity, wardrobe, palette, composition. Then say plainly what has to stay consistent across the output.

Strict editing

Open with "Strictly edit Video 1". State the one change, bound it, and list what is preserved. Never use the word reference in an edit prompt. It reads as a signal to generate something new instead of altering what you gave it.

Forward extension

Write "Extend Video 1 forward". The new footage picks up from the final pose, composition, motion direction, lighting, and audio of the source, then carries on.

Backward extension

Write "Extend Video 1 backward". You are building the lead-in, so the new footage has to land exactly on the state of the original first frame.

Track completion

Joining clips means spelling out the order of the sources and the transition between each pair. Three videos is the ceiling. Past that the model loses track of which clip it is on.

Combined tasks

Keep each authority in its own clause: "Reference the wardrobe from Image 1; strictly edit Video 1; replace the jacket." One instruction per clause, separated by semicolons.

Strict Edit Prompt

One bounded change, with an explicit preserve list.

Strictly edit Video 1. Change only the jacket colour, from olive to deep red.

Preserve: the person, the face, the hair, the camera movement, the framing, the background, the lighting, the pacing, and the full audio track.

Do not add or remove any subject. Do not restage the shot. No subtitles, no on-screen text, no watermark.

Motion & Continuity

Motion is where vague prompts fall apart. The model cannot animate a feeling, but it can animate a body part moving a certain distance at a certain speed. Write the second thing.

Break action into five parts

Body part, direction, amplitude, speed, force. "Raises the right arm shoulder high, slowly, with effort" gives the model far more to work with than "reaches up".

Describe the transition

Between two actions, say how the body gets from one to the other. Without it you get a snap. "Lowers the arm, shifts weight to the left foot, then turns" reads as one continuous move.

Show the emotion

Abstract states do not render. Replace "nervous" with the visible version: eyes flicking to the door, fingers tapping the cup, shoulders drawn in.

Order beats chronology

Chronological shot order carries more reliably than exact timestamps. Say what happens next, not what happens at 6.2 seconds.

Lock the cast

Never swap two characters, duplicate one, or change how many people are on screen mid-prompt. Identity and wardrobe stay fixed from the first shot to the last.

Chain shots on end state

A cut works when the next shot starts from the state the last one ended on. Write the end state, then open the following shot from it.

Failure prevention

Close long prompts with a short constraint line covering the five common failures: burned-in subtitles, logos, watermarks, a duplicated subject, and style drift between shots. Keep it to one line. Long exclusion lists get diluted, as covered in the next section.

Clean output: no captions, no stray music

Seedance has no negative prompt parameter on any endpoint. There is no separate field for things you want removed. Everything you want to suppress goes in the prompt itself, written as plain natural language.

End with a subtitle ban

Close your prompt with "No subtitles, no on-screen text." ByteDance recommends ending with "Keep it subtitle-free." Vertical 9:16 clips pick up spurious subtitles far more often than 16:9, so this matters most for vertical.

Keep ban lists short

One or two targeted exclusions work best. Long generic lists like "no text, no words, no letters, no logos, no watermarks" get diluted, and the model starts ignoring them.

Control the soundscape

Seedance generates music, dialogue, and sound effects together. Audio is all-or-nothing. For dialogue without a music bed, write "Audio: dialogue clear and prominent, ambient sound subtle. No music." The bare phrase "No music." tests more reliably than "no background music". Always describe the sound you do want.

Write dialogue as attribution

Use narrative attribution: Mara glances up. Mara says, quietly proud, "We finally did it." Keep lines under 10 words, one speaker per clip. Never ask for "voiceover" or "narration". Those words invite burned-in captions.

Dialogue Scene, No Music

Spoken dialogue over a quiet ambient bed, with no burned-in text.

A cramped mission control room at night, banks of monitors casting blue light across tired faces. Mara leans back from her console and glances up. Mara says, quietly proud, "We finally did it." Audio: dialogue clear and prominent, ambient sound subtle. No music. No subtitles, no on-screen text.

Clean Vertical Clip

A 9:16 vertical with sound effects only and every text ban in place.

9:16 vertical. A street food vendor flips noodles in a flaming wok at a night market, close-up on the tossing motion, steam and embers rising into neon light. Sounds of sizzling oil and clattering utensils. No music. No subtitles, no on-screen text. Keep it subtitle-free.

Frequently Asked Questions