Seedance 2.5.
50 references, 30 seconds, full edit control.
The complete prompting reference for Seedance 2.5: the core formula, reference-material roles, multi-reference scenes, 30-second stage structure, video editing, extension, keyframes, blockouts, transitions, and cinematography language. Every template is copy-ready.
Seedance 2.5 is live. Text-to-video, first and last frames, and multi-reference generation all run on the Seedance 2.5 page. Video editing and extension are a separate tool, because Seedance decides which of those you meant from the wording of your prompt.
Overview
Seedance 2.5 generates video from text alone, or from image, video, and audio references. It can also edit an existing video. Start by stating what you want to generate, then add reference materials, event progression, visual treatment, and audio as needed.
The headline changes in 2.5 are the reference budget (up to 50 materials in one generation), clips up to 30 seconds, scoped video editing, and forward and backward video extension. Every one of those depends on the same skill: telling the model which material does what, and what it must not do.
- Up to 50 reference materials across images, video, and audio.
- Clips up to 30 seconds, organised as stages with explicit end states.
- Video editing with a defined master, edit scope, and preserve list.
- Forward and backward extension anchored on the boundary frame.
- First and last frames, multi-keyframe sequences, storyboard grids, and blockout rendering.
- Dedicated syntax for music, sound effects, dialogue, and subtitles.
What each mode has to declare
Each mode has a small contract. Miss a term in the contract and the model fills the gap on its own. The sections below cover each mode in full, but this is the short version to check before you submit.
| Mode | The prompt must declare |
|---|---|
| Video editing | One video named as the sole editing master, the edit goal, the exact scope of the change, and the content that must survive untouched. |
| Forward extension | That the source clip’s last frame is inherited as the new first frame. State that before you describe any new action. |
| Backward extension | The action that happens before the source clip, then the source clip’s first frame as the final frame of the extension. |
| Storyboard or blockout | Which staging, motion, camera, and lighting traits to inherit, plus an explicit replacement for the blockout geometry, colors, and materials. |
| Multi-stage narrative | Named stages or time beats, each with an initial state, one primary event, and a visible end state. |
The Core Prompt Formula
Prompts can flexibly combine the following elements.
Subject + Action or Event + Scene and Environment (optional) + Visual Style (optional) + Camera Movement or Cut (optional) + Audio (optional)
- Subject + Action or Event: state who or what is doing what. This is the foundation of the video.
- Scene and Environment: describe the location, time, weather, spatial relationships, and background state.
- Visual Style: describe lighting, color, materials, image texture, or the overall mood.
- Camera Movement or Cut: describe shot size, camera angle, camera movement, the focus subject, and shot transitions.
- Audio: describe dialogue, voice characteristics, ambience, sound effects, and music.
<Subject> performs <primary action or event> in <scene and environment>. The visuals feature <visual style>. Use <shot size, camera angle, camera movement, or cuts>. Audio includes <dialogue, ambience, sound effects, or music>.
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Soft morning light enters through the window. The wet clay has a delicate sheen, and the workbench remains tidy. Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup's surface texture, then cut to a frontal view of the shelf. Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
Omit any component you do not need. Generation parameters do not belong in the prompt. Set those on the generation page or through the API.
Reference Materials and Their Roles
Material quantity and selection
Seedance 2.5 can combine up to 50 reference materials. Each material type follows the input limits below. The recommended ranges exist to improve generation stability. They are not capability limits.
| Material type | Input limit | Recommended range |
|---|---|---|
| Images | Up to 30 images, each no larger than 4K | Prefer 1 to 8 distinct subjects across subject-reference images |
| Videos | Up to 10 videos, combined duration no more than 30 seconds | Prefer 1 to 5 distinct subjects and 5 to 10 seconds per subject video |
| Audio | Up to 10 audio clips, combined duration no more than 30 seconds | Keep only dialogue, voice characteristics, ambience, or music directly relevant to the task |
| Video editing | A source video may be used together with reference images | Prefer a source video under 20 seconds and 1 to 5 reference images |
You can push past these ranges: 9 to 12 subjects in subject images, 6 to 10 subjects in subject audio or video, or 6 to 8 reference images for video editing. Stability drops as the material count grows. If more than five subjects also need multiple views, put each view in its own image. Independent view images are usually more stable than several views combined into one collage.
Token form
Reference tokens are short and lowercase, with no space inside them: @image1, @video1, @audio1. The spaced form used elsewhere in this guide points at the same slot. Pick one form and keep it consistent inside a single prompt so no token is read as two words.
One token carries exactly one production role. Do not let @image1 mean the lead character in one line and the set in another. If a material needs to supply two things, say both in the same line where the token is introduced.
Plan the registry before you write
With more than a handful of materials, sketch a registry first. It takes a minute and it catches the two mistakes that cost a generation: a token bound to nothing, and a token quietly bound to two different roles.
| Registry column | What to write |
|---|---|
| Token | The exact token as it will appear in the prompt, such as @image3. |
| Role | The one named subject, prop, scene, motion, or sound it stands for. |
| Inherit | The specific attributes to carry over, such as facial features, clothing, material, spatial layout, or pacing. |
| Exclude | Anything in that material that could leak into the output: background, bystanders, composition, on-image text, or the original color. |
Write an exclusion whenever leakage is plausible, not only when it is certain. A reference photo shot in a kitchen tends to bring the kitchen along unless you say otherwise. Once the registry is settled, every prompt line should agree with it. An instruction that contradicts a token’s declared role is the most common cause of a drifting subject.
Define each material’s role
After uploading reference materials, specify exactly what each one contributes. Add exclusions when people, backgrounds, or compositions in a material could be carried into the output unintentionally. Material mappings must be written in the prompt. Do not rely on text labels inside images, and do not make the model infer which person, prop, or scene each material represents.
@Image 1 defines <subject>'s <appearance, clothing, structure, or material>. @Video 1 defines <motion, camera movement, or pacing>. @Audio 1 defines <character or sound type>'s <voice, dialogue, ambience, or music>. <Subject> completes <primary action or event> in <scene>. The visuals feature <visual style>, with <camera treatment>.
@Image 1 defines the ceramic artist's facial features, hairstyle, and dark green apron. Do not use the image background. @Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image. @Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person's identity, clothing, or scene from the video. The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf. Begin with a medium shot of the wheel-throwing process, then slowly push in toward the cup's surface texture. Retain the sound of the wheel, the friction of clay, and indoor ambience.
Multiple views of one subject
If several images show different views of the same person or product, say so explicitly.
@Image 1 defines the front view of the same folding desk lamp. @Image 2 defines the left-side structure of the same folding desk lamp. @Image 3 defines the right-side structure of the same folding desk lamp. @Image 4 defines the rear structure of the same folding desk lamp. All four images define one folding desk lamp. The output must contain only one lamp throughout.
When a reference video already defines the motion, camera movement, and sequence accurately, state only which attributes to inherit. There is no need to restate every action, and repeating the motion description can conflict with the reference itself. A blockout video mainly supplies motion and spatial structure, so the prompt still has to define the intended subjects, scene, action, and visual style.
Audio and Text Syntax
Prompts can be written entirely in natural language. When you need to separate music, sound effects, dialogue, and subtitles more explicitly, use this syntax.
| Content | Syntax | Example |
|---|---|---|
| Music | ( ) | (Soft, rhythmic piano music plays in the background) |
| Sound effects | < > | <A bell rings in the distance> |
| Dialogue | { } | {Hello, welcome back.} |
| Subtitles | 【 】 | 【Chapter One: Departure】 |
Dialogue language reinforcement
When dialogue is not in Chinese, specify the language before the line. If the dialogue text is in English but the model speaks it in Chinese, or you need a specific regional variety, reinforce the language before the line.
Dialogue Language + Regional Variety or Accent + Delivery Style + Speaker + {Dialogue}
The girl says softly in Japanese: {もう大丈夫です}
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}Place audio inside the stage
In a staged prompt, keep each sound with the event that causes it instead of collecting all the audio in a closing paragraph. A lid click written into the stage where the lid is seated lands on the action. The same line written at the bottom of the prompt can drift anywhere in the clip.
The same applies to dialogue. Put the line in the stage where the character speaks, and state the language before it if it is not Chinese. Sounds that run under the whole clip, such as a music bed or room ambience, are the exception. Those belong in one line that covers the full duration.
Multi-Reference Creation
Seedance 2.5 supports up to 50 reference materials. With that many, the goal is not to cram every reference into one sentence. It is to define the relationships among characters, props, scenes, actions, and audio.
Define each material’s role → Map subjects → Group by type → Create subject profiles → Select references by scene
Step 1: name and map each subject individually
Bind each person, product, and prop to its reference material separately.
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Prop A> corresponds to @Image 3. Use only the structure, material, and color. <Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
Do not write "@Images 1 through 4 define four characters respectively." That wording never states which image belongs to which character.
Step 2: group materials by type
[Characters] <Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing. <Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing. <Exhibition Installer> corresponds to @Image 3. Use only the appearance, hairstyle, and clothing. <Guide> corresponds to @Image 4. Use only the appearance, hairstyle, and clothing. Do not interchange the four characters' appearances, clothing, actions, positions, or dialogue. [Props] <Sample Case> corresponds to @Image 5 and belongs only to <Conservator>. <Record Board> corresponds to @Image 6 and belongs only to <Registrar>. [Scenes] <Conservation Lab> references @Image 7. Use only the space, materials, and lighting. <Gallery> references @Image 8. Use only the space, materials, and lighting. [Motion and Audio] @Video 1 defines the motion of <Conservator> opening <Sample Case>. Do not use the person or scene from the video. @Audio 1 defines <Guide>'s voice and specified dialogue.
Step 3: create a centralized profile for important subjects
When the same character uses several references across multiple scenes, add a subject profile.
[Subject Profile: Conservator] Appearance and clothing: @Image 1. Fixed prop: <Sample Case> from @Image 5. Locations: <Conservation Lab> and <Gallery>. Motion references: the case-opening motion from @Video 1 and the sample-placement motion from @Video 2. Do not use: other characters' clothing. Do not give this character <Record Board> or guide equipment.
Step 4: select references by scene
Scene 1 | Inspection in the Conservation Lab Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the case-opening motion from @Video 1. Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside. End state: <Conservator> remains on the inner side of the workbench. <Sample Case> stays beside the conservator's right hand, which is on the left side of the frame. Scene 2 | Registration in the Gallery Use: <Registrar>, <Record Board>, and <Gallery>. Event: <Registrar> checks the number on <Record Board> beside the display case. End state: <Registrar> still holds <Record Board> with both hands. No other character enters the display-case area.
The goal of multi-reference creation is to help the model select the correct materials for the current scene. It is not to make every material appear at the same time.
30-Second Videos: Stages and End States
When a video contains several events, divide the story into consecutive stages. Give each stage only one primary state change, and state what should be directly visible at the end of that stage.
Pick the time structure first
Structure follows length and intent. Reaching for stages on a simple one-shot clip adds rigidity you do not need. Writing loose prose for a five-event story leaves the ordering to chance.
| Clip | Structure to use |
|---|---|
| Single shot, one event | Plain prose. State the subject, the action, and the scene. No stage headings. |
| 10 to 30 seconds, several events | Named stages. Each one gets an initial state, one primary event, and a visible end state. |
| Pacing-critical moments | Timestamp ranges, used only for entrances, exits, handoffs, and sound cues. Keep the ranges consecutive and non-overlapping. |
| Longer than 30 seconds | Coarse time blocks rather than second-by-second beats. Fine timing over a long clip tends to produce cuts you did not ask for. |
The two can be mixed. Run the story on named stages, then attach a timestamp to the one beat that has to land on time, such as the moment a second character enters or a door slam hits.
The end-state contract
A stage’s end state is a contract with the stage that follows. Whatever you leave out of it is free to change. So write down the details that carry continuity: who stands where, who is holding which prop, how many of each object are in frame, and which way people and objects face across the frame.
[Generation Goal] Generate a <video type>. The central subject is <subject>, and the primary event is <story summary>. [Stage 1] Initial state: <initial state of characters, props, and scene>. Primary event: <one primary action or event>. End state: <character positions, prop ownership, or visible scene state>. [Stage 2] Initial state: <continues from Stage 1's end state>. Primary event: <one primary action or event>. End state: <character positions, prop ownership, or visible scene state>. [Stage N] Initial state: <continues from the previous stage's end state>. Primary event: <one primary action or event>. End state: <final visible state>. Keep <character identities, clothing, prop ownership, and spatial relationships> consistent throughout.
Generate an instructional video showing a flower shop's order-packing process. <Florist> and <Store Assistant> arrange, wrap, and hand off a bouquet together. [Stage 1] Initial state: <Florist> stands behind the workbench. Loose flower stems, scissors, and wrapping paper lie on the tabletop. Primary event: <Florist> trims the stems and gathers them into a bouquet. End state: the finished bouquet rests in <Florist>'s left hand, and the scissors lie back on the tabletop. Keep <Florist> and <Store Assistant>'s identities, clothing, workbench orientation, scissors position, and bouquet ownership consistent.
Timestamps and pacing control
For ordinary narratives, use stages by default. Use one-second precision only when you need to control a critical handoff, an entrance or exit, a transition, or an explicit beat.
| Pattern | Example |
|---|---|
| Time range | 0-3 seconds... 3-7 seconds... 7-12 seconds... |
| Exact time point | At 5 seconds, the camera whip-pans rapidly to the left and completes the transition. |
| Relative timing | Three seconds after the character presses the button, the room lights gradually turn off. |
0-5 seconds: Show an empty wooden display table. A hand places a white ceramic plate on it. End state: the hand has left the frame, and only the white plate remains in the center of the table. 5-10 seconds: Remove the white plate, then place a clear glass on the table. End state: only the clear glass remains in the center of the table. 10-15 seconds: Remove the clear glass, then place a green ceramic vase on the table. End state: only the green vase remains in the center of the table.
Time ranges should be consecutive and non-overlapping. They represent an event’s time budget, not a precise edit point, so actions may land slightly before or after a boundary. Too little content in a range gives the model more freedom. Too much can cause excessive cutting or omitted events. Do not use timestamps to demand frequencies such as "complete three actions in one second".
Worked example: registry, stages, and one timed beat
This one puts the pieces together. The registry comes first, then three stages that hand state to each other, then a single timestamp on the beat that has to land on time. Audio cues sit inside the stages rather than in a separate paragraph, so each sound is tied to the event that causes it.
[Reference Registry]
@image1 defines <Barista>'s face, short curly hair, and black canvas apron. Do not use the image background or the second person in it.
@image2 defines <Courier>'s face and yellow rain jacket. Do not use the image background.
@image3 defines <Ribbed Travel Mug>'s shape, matte steel finish, and black lid. Do not use the image's studio backdrop or its reflections.
@image4 defines <Corner Cafe>'s counter layout, tiled wall, window position, and morning light. Do not use the people in the image.
@audio1 defines <Courier>'s voice. Dialogue language: American English.
[Generation Goal]
Generate a 20-second single-location scene. <Barista> prepares <Ribbed Travel Mug> and hands it to <Courier> across the counter of <Corner Cafe>.
[Stage 1]
Initial state: <Barista> stands behind the counter on the right of frame. <Ribbed Travel Mug> sits empty beside the machine. No other person is in frame.
Primary event: <Barista> fills <Ribbed Travel Mug> and seats the black lid.
End state: <Ribbed Travel Mug> is closed and stands on the counter near <Barista>'s right hand. <Barista> is still behind the counter, facing the camera. Exactly one mug is in frame.
Audio: <The espresso machine hisses, then a soft click as the lid seats>.
[Stage 2]
Initial state: continues from Stage 1's end state.
Primary event: <Courier> enters through the door on the left of frame and walks to the counter.
End state: <Courier> stands on the customer side of the counter, on the left of frame, facing <Barista>. <Ribbed Travel Mug> has not moved.
Audio: (a quiet acoustic bed continues under the scene). <Courier> says: {Order for the corner shop?}
[Stage 3]
Initial state: continues from Stage 2's end state.
Primary event: <Barista> slides <Ribbed Travel Mug> across the counter and <Courier> takes it with the right hand.
End state: <Ribbed Travel Mug> is in <Courier>'s right hand. <Barista>'s hands rest empty on the counter. Both characters keep their sides of the counter. Still exactly one mug in frame.
[Timed beat]
At 14 seconds, the mug passes from <Barista>'s hand to <Courier>'s hand in one continuous move, with no cut across the handoff.
Keep <Barista> and <Courier>'s identities and clothing, the mug count, the counter orientation, and the left-to-right screen positions consistent across all three stages.Read the example backwards to see why it holds. Every stage names who is on which side of the frame, which hand owns the mug, and how many mugs exist. Those three facts are the ones that drift first, so they appear in every end state.
Parameters That Lock Automatically
Video editing, first-frame or first-and-last-frame generation, and video extension automatically lock some generation parameters based on the input materials. Locked parameters cannot be set separately on the generation page or through the API.
| Task type | Aspect ratio | Duration |
|---|---|---|
| Video editing | Automatically preserves the input video’s aspect ratio. Cannot be set separately. | Automatically preserves approximately the input video’s duration. Cannot be set separately. Input-frame processing may introduce a difference of up to about 0.3 seconds. |
| First-frame or first-and-last-frame generation | Automatically uses the first image’s aspect ratio. The first and last images should share the same ratio to avoid stretching the last frame. | Can be set. |
| Video extension | Automatically preserves the input video’s aspect ratio. Cannot be set separately. | Can be set. |
Video Editing
When editing an existing video, first define the source video as the sole editing master. Then specify the edit target, edit scope, target material, and content to preserve. The output preserves the input video’s aspect ratio and approximately its duration. Input-frame processing may shift the duration by up to about 0.3 seconds, usually because of transition-frame handling, while the overall content and event order stay substantially unchanged.
General editing pattern
[Edit Goal] Edit @Video 1. Within <the entire video or a specific time range>, <add, remove, replace, or adjust> <visual object, region, or audio category>. [Source Video Role] @Video 1 is the sole editing master. It defines <characters, scene, actions, composition, camera movement, occlusion relationships, audio, and event order>. [Target Material Role] @Image 1 or @Audio 1 defines <specified attributes of the target object or sound>. [Edit Scope] Modify only <object, region, time range, or audio category>. [Content to Preserve] Keep <visual content, motion, audio, and timing relationships that must not change> from @Video 1.
Edit @Video 1. Only from 4-7 seconds, change the cool blue light on the right wall to warm orange light. @Video 1 is the sole editing master. It defines the character, room layout, actions, composition, camera movement, audio, and event order. Change only the light color on the right wall and the area it illuminates. Allow the character's skin tone to respond naturally to the environmental light. Keep the character's identity, clothing, expression, position, motion, room structure, camera movement, dialogue, and ambience from @Video 1.
Subject replacement
Edit @Video 1. Change only <original object> to <target object>. @Video 1 is the sole editing master. It defines the original scene, camera position, camera movement, motion path, occlusion relationships, and event order. [Target Reference Role] @Image 1 defines <target object>'s <appearance, structure, or material>. Do not use <irrelevant background, people, or composition>. Modify only <specific object and area>. The entire video contains <number> target object(s). Do not modify <content to preserve>. [Timeline Inheritance] <Target object> inherits every appearance, motion, occlusion, and exit of <original object>, including timing, duration, path, and speed changes. Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Edit @Video 1. Replace only the yellow folding desk lamp with the white folding desk lamp in @Image 1. @Video 1 is the sole editing master. It defines the desk, books, hand movements, camera position, camera movement, occlusion relationships, and event order. @Image 1 defines only the white folding desk lamp's appearance, structure, and material. Do not use the image's background, composition, or other objects. Keep exactly one white folding desk lamp throughout the video. Replace only the original yellow folding desk lamp. Do not modify the books, desk, hands, or background. The white folding desk lamp inherits every appearance, lamp-arm rotation, hand occlusion, and exit of the original yellow folding desk lamp, including timing, path, and speed changes.
Background replacement
Edit @Video 1. Replace only <original background area> with <target environment> from @Image 1. @Video 1 is the sole editing master. It defines the people, foreground objects, actions, composition, camera movement, and event order. @Image 1 defines only <target environment>'s spatial layout, materials, depth of field, ambient color, and lighting direction. Do not use the people or foreground objects in the image. Modify only <background outside the subject's silhouette>. Do not modify <subject identity, facial features, hairstyle, clothing, expression, position, size, or motion>. Keep the character actions and occlusion relationships from @Video 1. Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
@Video 1 is the sole editing master. It defines the people, actions, composition, camera treatment, and event order. @Image 1 provides only the spatial layout, depth of field, ambient color, and lighting direction of a daylit glass greenhouse. Do not use the people in the image. Replace only the light gray background outside the person's silhouette in @Video 1 with the daylit glass greenhouse from @Image 1. Keep the person's identity, facial features, hairstyle, clothing, expression, position, size, and arm-raising motion from @Video 1.
Audio editing
Dialogue, language, voice, background music, and sound effects can be edited separately. State the speaker or sound category, the intended change, and which other sounds must stay unchanged.
Edit @Video 1. Remove only the original background music. Keep the character dialogue, lip sync, ambience, and action sound effects; preserve the visuals, camera treatment, and editing rhythm from @Video 1. Edit @Video 1. Change <Presenter>'s spoken language to natural American English while preserving the dialogue content and speaking times. Keep all other character voices, background music, ambience, and visuals from @Video 1.
Video Extension
Video extension creates content beyond the boundary of a source video. For a forward extension, the extension’s first frame continues from the source video’s last frame. For a backward extension, the extension’s last frame connects to the source video’s first frame. Beyond the boundary frame, check that the characters, props, background, and events in the extended segment are correct.
Forward extension (after the original video)
First describe the continuous state of the source video’s last frame, then describe what happens afterward.
@Video 1 is the source video to extend forward. Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain continuity in <subject pose and orientation>, <prop position>, <background and spatial relationships>, <camera position and composition>, <lighting>, and <motion direction>. Then, <describe the new action, event, camera treatment, or audio to add>. Throughout the extension, maintain continuity in <character identity and clothing>, <key props>, <background layout>, and <axis of action>. Keep each subject as the same continuous instance throughout: do not duplicate or split it, and keep the person's appearance or the object's number of parts stable.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the same locked-off medium shot, the orange paper airplane's position and orientation, the classroom-window background, the afternoon lighting, and its movement toward the right side of the frame. Then, the orange paper airplane continues gliding toward the right and exits the frame while the white curtain beside the window sways slightly. Keep the camera and classroom background in the state established by the source video's last frame.
Forward extension with additional references
Define the role of every additional material first, then state that the source video controls the extension boundary. New materials may supplement characters, props, or audio, but they must not override the source video’s last-frame control over the extension’s opening image.
@Image 1 defines <Gardener>'s facial features. @Image 2 defines <Gardener>'s light green work apron. @Image 3 defines <Wicker Flower Basket>'s structure and material. Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the greenhouse workbench, <Gardener>'s position, and <Wicker Flower Basket>'s position. Then, <Gardener> lifts <Wicker Flower Basket> with both hands and places it on the middle shelf of the wooden rack behind them. Throughout the extension, maintain continuity in <Gardener>'s face, apron, greenhouse layout, and camera direction.
Backward extension (before the original video)
First describe what happens before the source video begins, then define the source video’s first frame as the explicit end state of the extended segment. Writing only "then connect to the source video" may introduce later characters or effects too early, or cause the image to keep changing after it reaches the target state.
@Video 1 is the source video to extend backward. Extend @Video 1 backward. Before the source video begins, <describe the preceding action, event, camera treatment, or audio>. The last frame of the extended segment naturally connects to the first frame of @Video 1: <subject pose and orientation>, <prop position>, and <background and spatial relationships>. Match the <camera position and composition>, <lighting>, and <motion direction> of @Video 1's first frame. <Materials that should appear only after the source video begins> must not appear early in the backward extension.
Extend @Video 1 backward. Before the source video begins, show an empty establishing shot of the same glass greenhouse. Morning mist drifts slowly near the floor, the overhead shade rises gradually, and no people are present yet. The last frame of the extended segment naturally connects to the first frame of @Video 1. Match the greenhouse's central aisle, planting tables on both sides, glass frame, soft morning light, and locked-off wide composition. At the end, the shade is fully raised, the aisle is empty, and the leaves still sway slightly.
Boundary frames should connect naturally at a visual level. That does not mean they will be pixel-identical. During review, inspect both sides of the boundary and the complete extended segment.
Keyframes, Storyboards, and Blockouts
First and last frames with additional references
In multimodal reference mode, state in the first line that @Image 1 is the first frame and @Image 2 is the last frame. There is no need to switch to a separate first/last-frame mode. The output aspect ratio locks to the first image, and duration is set on the generation page or through the API. The first and last images should use the same aspect ratio, since mismatched ratios may stretch the last frame. Additional images can still define characters, props, scenes, and materials.
@Image 1 is the first frame. It defines the opening composition, subject position, pose, prop state, scene, and camera direction. @Image 2 is the last frame. It defines the ending composition, subject position, pose, prop state, scene, and camera direction. @Image 3 defines <Subject A>'s <appearance, clothing, structure, or material>. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2. @Image 4 defines <specified attributes> of <Subject B, prop, or scene>. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2. <Describe one continuous action or event>. The video begins naturally from the first frame defined by @Image 1 and reaches the last frame defined by @Image 2 after the continuous action. Between the first and last frames, maintain continuity in <character identity, prop structure and ownership, scene layout, and camera direction>.
@Image 1 is the first frame. It defines the opening composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction. @Image 2 is the last frame. It defines the ending composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction. @Image 3 defines <Perfumer>'s face, hairstyle, and dark green apron. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2. @Image 4 defines <Glass Perfume Bottle>'s shape, material, and label position. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2. Starting from the first-frame pose, <Perfumer> picks up a dropper and <Glass Perfume Bottle>, drips amber fragrance oil into the bottle, swirls it gently, closes the stopper, places the finished bottle in the center of the table, and naturally reaches the last frame defined by @Image 2. Between the first and last frames, maintain continuity in <Perfumer>'s identity and clothing, bottle count and structure, wooden-table layout, warm side lighting, and camera direction.
Describe each anchor image separately. Do not combine them into a sentence such as "@Images 1 and 2 are the first and last frames." Other references should supplement only their specified attributes and must not replace the first-frame or last-frame composition.
Multi-keyframe sequence control
When separate images define different stages of a process, begin with "Use @Image 1 through @Image N as keyframes in this order", then describe the key state each image represents. Independent keyframe images are usually easier to align than several frames combined into one grid. They control stage order and key states. They do not reproduce every frame exactly.
Use @Image 1 through @Image N as keyframes in this order. @Image 1 is the first frame. It defines <opening composition, subject position, pose, prop state, and camera direction>. @Image 2 defines the second keyframe: <visible end state of Stage 1>. @Image 3 defines the third keyframe: <visible end state of Stage 2>. @Image N is the last frame. It defines <ending composition, subject position, pose, prop state, and camera direction>. The video passes through the states defined by @Image 1, @Image 2, @Image 3, and @Image N in order, using continuous action to transition naturally between stages. Maintain continuity in <subject identity, prop structure and ownership, scene layout, lighting, and axis of action> throughout.
Use @Image 1 through @Image 4 as keyframes in this order. @Image 1 is the first frame. It shows an orange paper airplane resting on the left side of a classroom desk, pointed toward the right, in a locked-off medium shot. @Image 2 defines the second keyframe: one hand lifts the same orange paper airplane from the desk without changing its direction. @Image 3 defines the third keyframe: the same orange paper airplane passes the window while the curtain moves slightly to the right. @Image 4 is the last frame. It shows the same orange paper airplane resting on the middle shelf of the bookcase on the right, still pointed toward the right. The video passes through the states defined by @Image 1, @Image 2, @Image 3, and @Image 4 in order. Keep flight direction and speed continuous between stages. Maintain the paper airplane's orange material, size, and folds, as well as the classroom layout, afternoon side lighting, and camera axis.
Storyboard grids
A storyboard grid communicates the overall story, shot order, and approximate compositions. It is not meant for strict reproduction of every detail in every panel. Prefer no more than 15 panels, use clean line art or simple diagrams, and minimise text labels. State the reading order, then describe each panel’s subject action, shot size or camera movement, final visual style, and audio.
@Image 1 provides an <N-panel storyboard grid> for shot order and approximate composition. Read it <left to right, top to bottom>. Do not use the grid's <line-art style, text labels, or placeholder characters>. @Image 2 defines <Subject A>'s <appearance and clothing>. @Image 3 defines <key prop or scene>'s <structure, material, or lighting>. Shot 1: <shot size, subject action, and scene state>. Shot 2: <shot size, subject action, camera movement, or transition>. ... Shot N: <closing action and final visible state>. The final video uses <visual style>. Audio includes <dialogue, ambience, action sound effects, or music>.
@Image 1 provides a four-panel pottery-making storyboard for shot order and approximate composition. Read it left to right, top to bottom. Do not use the storyboard's line-art style or text labels. @Image 2 defines <Ceramic Artist>'s face, short hair, and dark gray apron. @Image 3 defines <Blue-Glazed Cup>'s proportions, glaze color, and curved handle. Shot 1: a wide shot establishes a quiet pottery studio with <Ceramic Artist> seated at the wheel. Shot 2: a side medium shot shows both hands shaping the rotating clay as the cup body takes form. Shot 3: a close-up shows fingers refining the rim and handle joint while slip moves slowly over the fingertips. Shot 4: a medium close-up shows the fired <Blue-Glazed Cup> placed on a wooden shelf as <Ceramic Artist> withdraws both hands. Use a realistic documentary look. Retain the wheel's rotation, wet-clay friction, and studio ambience.
Coarse blockouts
Blockout references fall into two categories. Coarse blockouts mainly provide temporal information: action, paths, blocking, camera movement, cuts, lighting, and sound. Fine blockouts already contain complete structures and are mainly used for re-rendering materials, colors, characters, scenes, and visual style. Work out which one you have before choosing a prompt structure.
Use a coarse blockout to lock action paths, motion direction, blocking, entrances and exits, camera paths, cut points, lighting changes, and sound rhythm. Map each geometric object separately to its final subject or prop. Prefer simple geometry with clear relationships. Arms, wings, and other appendages should be used only when the action sequence is complete, otherwise they can cause stiff motion or structural misreading.
@Video 1 is a coarse blockout reference. It provides only <motion paths, subject blocking, camera position, camera movement, cuts, lighting changes, sound rhythm, or spatial relationships>. Do not use its blockout appearance, materials, or scene. <Blockout Subject A> in @Video 1 corresponds to <Subject A>. <Blockout Subject B or geometric prop> in @Video 1 corresponds to <Subject B or key prop>. @Image 1 defines <Subject A>'s <appearance, clothing, or structure>. @Image 2 defines <specified attributes> of <Subject B, key prop, or scene>. Keep <motion path, blocking, camera movement, cuts, lighting, or sound rhythm> from @Video 1. The final video uses <characters, scene, materials, and visual style>. Audio includes <dialogue, ambience, or action sound effects>.
@Video 1 is a coarse blockout reference. It provides only the character's walking path, cart direction, locked-off camera, one push-in, and two cuts. Do not use its gray geometry or empty scene. The tall cylinder in @Video 1 corresponds to <Guide>. The rectangular block in @Video 1 corresponds to <Mobile Display Cart>. @Image 1 defines <Guide>'s face, blue uniform, and name badge. @Image 2 defines <Mobile Display Cart>'s white metal frame and clear cover. @Image 3 defines the technology gallery's curved walls, gray floor, and overhead strip lights. <Guide> pushes <Mobile Display Cart> along the curved wall, stops in front of the central display, and opens the clear cover. Keep the walking path, subject blocking, push-in direction, and cut points from @Video 1. Use a bright, realistic museum-documentary style. Retain footsteps, wheel sounds, and gallery ambience.
Fine blockouts
A fine blockout already contains complete character, prop, or scene structures. Use it to change materials, colors, character appearance, scene, or overall visual style. Keep the blockout clean: remove path lines, coordinate axes, controllers, camera frustums, and other production markers.
@Video 1 is a fine blockout reference. Preserve <subject structure, action, spatial layout, camera position, camera movement, and cuts>. Do not use its original gray materials or empty background. @Image 1 defines <subject>'s <character appearance, material, color, or surface details>. @Image 2 defines <scene>'s <space, materials, lighting, or visual style>. Re-render <subject> from @Video 1 as <final subject>, and re-render the scene as <final scene>. Keep <structure, action, camera treatment, and spatial relationships> from @Video 1. Use <materials, colors, and style>. Audio includes <ambience, sound effects, or music>.
@Video 1 is a fine blockout reference. Preserve the kinetic sculpture's complete structure, three-ring rotation relationship, pedestal position, orbiting camera movement, and cuts. Do not use the gray materials or empty background. @Image 1 defines the outer ring's brushed-brass material. @Image 2 defines the inner blades' translucent blue-glass material. @Image 3 defines a contemporary gallery with white curved walls, a dark gray floor, and soft overhead lighting. Re-render the ring structure from @Video 1 as a kinetic sculpture made of brass and blue glass, and re-render the scene as a contemporary art gallery. Keep the structure, rotation rhythm, orbiting camera movement, and cuts from @Video 1. Retain the sculpture's low mechanical rotation sound and quiet interior ambience.
One-Click Video
One-click video organises multiple images, or images plus a style-reference video, into a complete video with consistent pacing and visual packaging. State each material’s role, image order, amount of motion, editing rhythm, visual treatment, and audio. Do not write only "turn these materials into a video".
Material Roles → Image Order → Motion Amount → Editing Style → Visual Treatment → Audio
[Material Roles] @Image 1 to @Image N provide <scenes, subjects, or products>. @Video 1 provides <editing rhythm, camera treatment, or visual style> only. Do not use its subjects or scene. [Image Order] Use the images in <the given order, or grouped by theme>. [Motion Amount] Apply <slight, moderate, or strong> motion to each image. [Final Style] Use <editing rhythm, transition method, and visual treatment>. [Audio] Include <dialogue, ambience, sound effects, or music>.
[Final Style] Use an upbeat travel-video rhythm. Connect scenes with natural occlusion and similar colors. Keep hand-drawn stickers at the frame edges. [Audio] Retain night-market chatter, light dish sounds, and riverside wind, with upbeat but unobtrusive instrumental music.
If image order matters, state the exact sequence. If the model may arrange the materials freely, say it can organise them by theme. When several characters or products appear, keep naming and binding each one separately.
Seamless Video Transitions
A seamless transition generates continuous bridge content between two videos. Identify the before-transition and after-transition clips first, then describe the trigger action, camera movement, visual transformation, arrival state, and audio transition.
Before Video → After Video → Trigger Action → Camera Movement → Visual Transformation → Arrival State → Audio
| Transition method | What to specify |
|---|---|
| Dive or reverse movement | Camera direction, speed change, and when the next scene begins |
| Character rotation | Pose, rotation direction, and how clothing or background changes continuously |
| Foreground occlusion | When the foreground object fills the frame and the composition that follows |
| Object morph | Corresponding shapes, materials, and the transformation process |
| Push/pull or focus change | Camera movement, focus target, and continuous spatial relationship |
@Video 1 is the before-transition clip. Use its <ending subject, action, composition, camera direction, and audio>. @Video 2 is the after-transition clip. Use its <opening subject, composition, camera direction, and audio>. Keep <character identity, product structure, scene, and primary action> stable in the original portions of @Video 1 and @Video 2. At the end of @Video 1, <subject or foreground object> triggers the transition through <action>. The camera <movement direction and speed change>, while <shape, material, light, or space> gradually transforms into <corresponding element> at the start of @Video 2. The transition ends naturally at @Video 2's opening composition, preserving continuity in <subject position, camera direction, and motion trend>. Audio transitions smoothly from <before audio> to <after audio>.
@Video 1 is the before-transition clip. Use its rainy night street, red umbrella, slow push-in, and rain sound. @Video 2 is the after-transition clip. Use its circular gallery skylight, upward camera movement, and quiet interior reverberation. Keep the people, street, gallery structure, and primary actions in the two original videos stable. At the end of @Video 1, the red umbrella approaches the camera and gradually fills the entire frame, triggering the transition. The camera continues moving forward. The umbrella's circular edge gradually becomes the skylight's metal ring, and the red fabric transitions into white daylight passing through the skylight. The transition ends naturally at @Video 2's upward-looking opening composition, with the camera movement changing smoothly from forward motion to an upward rise. The rain gradually fades into footsteps reverberating inside the gallery.
The goal of a seamless transition is visual and audio continuity. A prompt may ask to preserve the primary content of both source videos, but a generated bridge is not a pixel-identical edit splice.
Emotional Direction and Observable Performance
Emotion words such as "tense", "warm", or "oppressive" communicate an overall direction, but they leave a lot of room for interpretation in the performance. For more stable acting, add directly visible or audible cues: eye movement, brow tension, mouth movement, breathing, gaze direction, and hand movement.
You do not need to list every facial detail. For a single emotional transition, two to four clear cues are usually enough. Use event-triggered stages only when the emotion changes several times.
The overall emotion shifts from <starting emotion> to <ending emotion>. After <triggering event>, <subject> first shows <immediate observable reaction>. Then, <eyes, brows, mouth, breathing, gaze, or hand movement> gradually <changes>. Finally, <subject> expresses <target emotion> through <restrained or explicit outward behavior>.
After confirming <critical information>, the emotion that <subject> tries to restrain or conceal gradually becomes visible through <observable behavior>. Finally, <subject's final action, expression, or manner of speaking>.
Applause marking the end of the performance comes from behind the stage. The young actor's fingers suddenly stop on the program, the gaze turns slowly toward the curtain, and the shoulders remain tense. After confirming that the curtain call is over, the actor exhales softly. The shoulders gradually relax, a restrained smile appears, and the eyes slowly well with tears, but the actor never turns to leave.
Professional Cinematography Terms
Basic camera language and popular camera techniques can be written straight into the prompt. When a term is uncommon, has several interpretations, or needs precise control, also state which subject it applies to, how the image changes, and the intended visible result.
Basic camera language
| Type | Commonly supported terms |
|---|---|
| Shot size | extreme wide shot, wide shot, medium shot, close-up, extreme close-up |
| Camera movement | push in, pull out, pan, lateral move, follow shot, orbit, dive, dolly out, tilt up, handheld shake |
| Camera position and viewpoint | low angle, overhead view, first-person view |
Popular camera techniques
These can be used directly. If the frame contains several subjects, still say which subject the camera follows or revolves around, where the movement begins, and where it ends.
| Technique | What to specify |
|---|---|
| One-take shot | The subjects, spaces, and events the continuous camera passes through in order |
| Dolly zoom | The subject size to preserve and whether the background appears to move closer or farther away |
| Aerial view | Viewing height, movement direction, and the environmental area to reveal |
| FPV | First-person flight or traversal path, speed, and turns |
| Bullet time | The action to freeze or slow down and the camera's orbit direction |
| Handheld camera | The subject being followed and the amount of shake |
| Bounce speed ramp | Where the action accelerates, decelerates, or rebounds, and its final resting state |
Uncommon cinematography terms
For a niche term, a term with inconsistent industry usage, or a term the model may not recognise, keep the term itself and translate it into a directly observable visual change.
Cinematography Term + Target Subject + Visual Change + Foreground/Background Relationship + Direction or Speed
Rack focus: shift focus smoothly from the leaves in the foreground to the person in the background. The leaves gradually blur while the person's face changes from soft to sharp.
For a precise transition, also state the trigger time, occluding object, camera direction, transition method, and the composition or motion trend that should continue afterward.
Example 1. Shallow-depth-of-field portrait: keep <Pastry Chef>'s eyes and face sharp while the glass jars and lights in the background become soft, circular bokeh. Example 2. Tracking shot: move horizontally at the same speed as <Skateboarder>, keeping the subject sharp while the roadside wall forms horizontal motion blur from right to left. Example 3. Golden hour: warm, low-angle sunlight enters from behind and to the left of <Hiker>, casting long shadows across the mountain ridge. Example 4. Natural vignette: darken the four corners gradually while keeping the brightness and skin tone of <Pianist> in the center natural, without a black border. Example 5. Whip-pan transition: at 5 seconds, move the camera rapidly to the left. Cut when the foreground bookshelf fully covers the frame, then continue moving left at a similar speed in the next scene.
Aperture, focal length, and shutter values can be included, but the intended visible result is usually clearer than a numeric value on its own.
Pre-Submission Checklist
- Does the prompt clearly state the subject and primary action or event?
- Does every reference material state what to use and what not to use?
- Is every distinct character, product, and prop named and bound to a reference?
- Are references selected by scene instead of being required to appear all at once?
- Does each stage of a long video contain only one primary change and a clear end state?
- Do the number of characters, clothing, prop ownership, and spatial relationships stay consistent?
- For video editing, does the prompt define the sole editing master, edit scope, target quantity, and content to preserve?
- Are abstract emotions and cinematography terms paired with directly visible or audible cues?
- Are first and last frames and multiple keyframes assigned one role per image, and do the first and last images use the same aspect ratio?
- Does the storyboard state which structure to inherit? For blockouts, did you identify whether the reference is coarse or fine, and specify the temporal, structural, material, and style information to inherit?
- Do video editing, first/last-frame generation, and video extension follow their automatically locked aspect-ratio and duration rules?
- For video extension, did you check the boundary image, motion trend, and audio continuity?
- For one-click video, does the prompt define material roles, image order, motion amount, editing style, and audio?
- Does every token carry one role only, and does every later instruction agree with the role you gave it?
- Are the time ranges consecutive, non-overlapping, and used only where pacing actually matters?
- Do object counts hold across stages, so a single prop cannot become two?
- Is each sound or line of dialogue placed in the stage where it happens, apart from ambience or music that runs throughout?
- For an extension, is the seam fully described on both sides, including pose, prop position, composition, lighting, and motion direction?