Skip to main content
MiniMax H3 Prompting Guide

MiniMax H3.
2K, sound included, tested prompt by prompt.

A prompting guide for MiniMax H3 built from generations on the live model. Timed beats, on-screen text that actually renders, reference images, first and last frames, camera refusals, and directed sound. Includes the two techniques we could not show made any difference.

Overview

MiniMax H3 generates video from a text prompt, from a first frame and an optional last frame, or from up to nine reference images. Everything comes out at 2K and 24 frames per second with a stereo audio track in the same file. Clips run from 5 to 15 seconds, and the prompt field takes up to 7,000 characters, so a full shot list with sound design fits in one request.

This guide is written from generations on the live model rather than from a spec sheet. Every clip below is one of those tests, at the resolution it came out of the model. Where a technique earns its words, there is a template for it. Where a widely repeated technique made no measurable difference in our tests, that is what we say.

Text asked for by name, spelled out in the prompt

The prompt asked for a title card reading "CHAPTER ONE" in condensed white all-caps. That is what came back, correctly spelled and cleanly set, in a single generation.

  • 5 to 15 seconds in one generation, at 2K and 24fps.
  • Native stereo audio in the same file: room tone, foley, music, and speech.
  • Six aspect ratios from 21:9 through 9:16, all at 2K.
  • Up to nine reference images, or a first frame and a last frame. Not both.
  • Prompts up to 7,000 characters. 3 credits per second of output.

Mode is derived from what you attach, not from a dropdown. No media means text-to-video. A start frame (optionally with an end frame) means image-to-video. Anything you are treating as a reference rather than a literal frame means reference-to-video.

What Transfers From Other Models

Habits built on Seedance and WAN mostly carry over, and two of them turned out to be doing less work here than expected. Each row below was tested as a matched pair: same subject, same seed, same duration, with only the technique under test changed.

TechniqueResult on MiniMax H3What we saw
Spelling out on-screen textWorks, and text rendering is a genuine strengthBoth arms of the test produced a clean, correctly spelled title card. Spelling the string out did not make text appear, it decided which words appeared and how they were set: the described version invented "CHAPTER I" in a serif face, the spelled-out version returned "CHAPTER ONE" in the condensed all-caps that was asked for.
Subject, scene, action, camera, audio orderingWorks, buys control rather than qualityA bare four-word prompt still produced a competent clip, but the model wrote its own film: two setups and a cut. The structured version held one continuous shot and executed the single push-in that was asked for.
Timed beats in bracketsWorks, and end states are what they buyAt 10 seconds the plain version already sequenced all three actions correctly. The timed version additionally hit the stated end states, including the empty bench after beat one.
Translating a camera instruction into an explicit refusalWorks, and is worth the wordsSaying nothing about the camera produced continuous drift and reframing. "The frame never moves", with the moves it should not make listed, held one fixed frame for the whole clip.
Observable behaviour instead of emotion wordsWorksNamed physical cues (fixed downward gaze, still fingers, lifted shoulders) were followed literally and held for the whole clip. "She looks anxious and nervous" produced a generic to-camera performance instead.
First and last frameWorks, and lands preciselyGiven both frames, the clip started on the supplied opening frame and resolved exactly onto the supplied closing frame, down to the coiled rope on the empty dock.
Assigning a job to every reference imageNo measurable difference in our testsTwo pairs, including a deliberately hard one with a person, a product on white, and an empty room. The arm with no role language composited exactly as well: right face, right bag, right room, and the white product background excluded without being told.
Negative listsCould not be shown to do anythingTwo pairs. In both, the control refused to produce the banned elements anyway, so there was nothing for the negatives to suppress. One small difference: the control rendered lettering on a shopfront while the arm that banned on-screen text did not.

The short version: structure, timing, camera refusals, observable behaviour, and spelled-out text all earn their words. Reference roles and negative lists did not change anything we could measure, which is not the same as proving they never do.

A note on the two negative results, because they contradict guidance published elsewhere. Our reference tests used inputs that could only fill one slot each, which is the easy case. Role language most plausibly matters when two references could plausibly be the same thing, such as two people or two locations, and we did not test that. Likewise, our negative lists banned things the model was not inclined to add. Both techniques are free, so there is no reason to drop them. They are just not the load-bearing part of a prompt here.

The Core Prompt Formula

Subject + Action + Scene and Environment + Visual Style + Camera + Audio

H3 fills gaps confidently. Leave the camera unspecified and it will direct itself, often cutting between two setups inside five seconds. That is a good deal when you want a finished-looking clip and a bad one when you need a specific shot. Name the parts you care about and let it invent the rest.

Basic template
<Subject> <performs the primary action> in <scene and environment>, at <time of day>.
<Light direction, materials, colour, and texture>.
Begin on <opening shot size and framing>, then <the one camera move you want>.
Audio: <the sounds that belong to this action, in the order they happen>.
Example, tested
A barista in a narrow espresso bar tamps a portafilter, locks it into the group head, and starts the shot, at sunrise.
Low warm sunlight comes through a window on the left, steam catches the light, brushed steel and matte ceramic surfaces.
Begin on a medium shot of her hands at the machine, then slowly push in to a close-up of the espresso stream.
Audio: the grind of the burr, the clack of the portafilter locking in, the hiss of the pump, and quiet morning room tone.
Bare: "A barista making coffee."

Competent, and entirely the model’s own film. It picked a macro latte-art pour, then cut to a wide of a barista it invented.

Structured: the template above

One continuous take. The push-in that was asked for, light from the stated side, ending on the espresso stream. Same seed as the bare version.

Neither clip is bad, and that is the finding. The structured prompt did not buy prettier pictures, it bought obedience. Use the formula when the shot has to match something already in your head, and a one-liner when you are happy to be surprised.

Timed Beats and Longer Clips

For anything longer than a single action, write the clip as consecutive timed beats. Give each beat one primary change and say what should be visible when it ends. The end state is the part that does the work: it gives the model a checkpoint to hit rather than a mood to sustain.

Timed-beat template
<One line naming the setting and the overall event.>
[0-3 seconds] <Opening shot. One primary change.> End with <a state a viewer could point at>.
[3-7 seconds] <Next shot. One primary change.> End with <a state a viewer could point at>.
[7-10 seconds] <Final shot. One primary change.> End with <the closing state>.
Audio: <what runs underneath, and what happens at the specific moments>.
Ten seconds, three actions in one sentence

The sequence still plays in the right order. H3 is good at implicit sequencing, so a plain description is not the disaster it is on some models.

Ten seconds, the same actions as timed beats

Watch the stated end states land: the bench is visibly empty once the wheel is in the stand, and the wrench is still in his hand at the end of the spoke work.

Where fifteen seconds runs out

We ran the 15-second ceiling with five beats of three seconds each. The first four landed in order and on time. The fifth, which required cutting the base free with a wire and lifting the finished vase away, did not complete before the clip ended.

Fifteen seconds, five timed beats

Centring, opening, pulling the walls up, shaping the lip. The final beat, lifting the vase off an empty wheel head, gets compressed and never finishes.

  • Budget about four seconds for any beat that involves a prop change or a hand-off. Three is enough for a camera change but tight for an action.
  • Expect the last beat to be the one that gets squeezed. Put the shot you care about most in the middle, not at the end.
  • One primary change per beat. Two changes in one beat collapse into whichever is easier to render.
  • Write end states as things a viewer could point at: an empty surface, a tool in a named hand, a door that is now closed.

On-Screen Text

Legible text is the thing H3 is most obviously good at, and it is the biggest practical difference from the other video models we have tested. Both halves of our test produced clean, correctly spelled type. What the prompt controls is which words appear and how they are set.

Described: "a title card with the name of the chapter"

Clean and correctly spelled, but the model chose the wording and the face: "CHAPTER I", set in a serif.

Spelled out: reads "CHAPTER ONE", condensed white all-caps

The exact string, in the requested weight and case, centred. Same seed as the described version.

Text template
A title card reads "<THE EXACT STRING>" in <weight, case, and family>, <position in frame>.
Do not misspell it, do not add any other text, do not add subtitles.
  • If a word has to be readable, type the word. Describing it gets you text, but not your text.
  • Name the typographic treatment: condensed, all-caps, serif, tracked wide. It follows these.
  • Say where it sits in frame. "Centred" and "lower third" both land.
  • Add the do-not line. It costs nothing and it is the one place a negative earned its space in our tests.

Reference Images

Attach up to nine images and H3 composites them into one shot: a face from one, a product from another, a location from a third. Published guidance says the highest-leverage habit is telling the model what each image is for. We tested that twice and could not measure a difference.

Three references, every role assigned

Image 1 the woman, Image 2 the satchel, Image 3 the lobby, plus explicit instructions not to use the white product background or the studio backdrop.

The same three references, no role language

Same seed, same images, one plain sentence. Right face, right bag, right room, and both unwanted backgrounds dropped without being told.

The earlier two-image test with a person and a location came out the same way. Our reading is that when each reference can only plausibly fill one slot, H3 works out the assignment on its own. Role language is free and we still recommend writing it, because the moment two references could fill the same slot the model has to guess, and a sentence is cheaper than a re-run.

Reference template
Use Image 1 for <what it defines>: preserve <the specific features that must survive>.
Use Image 2 for <what it defines>: preserve <the specific features that must survive>.
Use Image 3 for <the location>: <the elements of the set that matter>.
Do not use <the parts of any reference that should not appear, such as a studio backdrop>.
<The action, the light, and the shot.>

One thing worth knowing: faces and hair held across every reference test, but wardrobe drifted. A navy canvas jacket came back as denim in both arms. Name the garment in the prompt as well as showing it, and check it before you build a sequence around it.

On PixelDojo the references are images. The underlying model also reads reference video and reference audio for motion transfer and voice, which the tool does not expose today.

Advanced Reference Prompting

Everything above is enough for a single shot with one or two references. Once a prompt carries several assets and a named cast, it needs a token system and a fixed layout. This section is the structure we use for those, written from the model documentation rather than from our own A/B battery.

Reference tokens

Refer to attached assets with compact tokens: Image1, Video1, Audio1. No space before the number. The number is the upload position, so a token means whatever sat in that slot when you attached it. Fix the order before you start writing and never renumber mid-prompt. If you swap an asset, swap the file rather than the token.

  • Map every character and every voice one at a time. "The two women, respectively" is the fastest way to get them crossed.
  • Use the token every time you mean that asset. Falling back to "the reference photo" halfway down a long prompt reopens the guess.
  • One job per asset. If a single image has to supply both a face and a room, say which part supplies which.
  • Say what to exclude as well as what to keep. A studio backdrop, a white product sweep, or a watermark is part of the file and needs ruling out by name.

Pick the task contract first

Before writing a word, decide which of the narrow contracts the shot is. Each one has a different rule about what the references are allowed to control. Mixing two contracts in one prompt is where most reference work falls apart.

ContractWhat the references controlWhat the prompt must say
Text to videoNothing is attachedThe whole shot, in the core formula above.
First and last frameThe two boundary imagesWhich image opens and which closes, plus one causal motion that gets from one to the other.
Omni referenceIdentity, wardrobe, props, and locationOne named job per image, and the features that must survive the composite.
Performance transferOne source video supplies motion onlyThe named motion to copy, and an explicit refusal of that video’s people, clothing, and location.
Voice transferOne audio file supplies a voiceWhich character speaks in that voice, that lip sync holds, and what the room sounds like underneath.
Targeted editOne clip is the masterThat the master is the sole source, the exact list of changes, and the list of things that must not move.

The rule underneath all six: an asset controls one thing, and everything it should not control is written down. "Use the references" is not an instruction, it is a hope.

Lock the stable truths before the shot list

Anything that has to stay identical for the whole clip belongs above the shot list, in labelled sections. The shot list then only carries what changes. Subject count, faces, wardrobe, props, voices, and who stands where are the six that drift if you leave them to the beats.

Section layout
[REFERENCE USE]
Image1 = <one job>. Keep <features>. Ignore <what must not travel>.
Image2 = <one job>. Keep <features>. Ignore <what must not travel>.

[IDENTITY / CONTINUITY LOCKS]
<How many people are in frame, and no more.>
<Each face, tied to its token.>
<Each garment, described in words as well as shown.>
<Each hero prop, and which hand or surface holds it.>

[SCENE]
<Location, time of day, weather, and the light.>

[DIALOGUE]
<Character>: "<the exact line>"

[SCREEN GEOGRAPHY]
<Who is on the left, who is on the right, and what they look at.>

Screen geography is the one people skip. Naming who stands where, and where each pair of eyes goes, is what keeps eyelines from breaking when the camera moves.

Then the shot list, then the finish

The shot list follows the same timed-beat rules as earlier in this guide: consecutive ranges that do not overlap, one primary event each, and a stated end state. Give each range a framing, a camera behaviour, an action, any line of dialogue, and what is true when it ends. For a single continuous action, plain prose is still fine. Brackets are for sequences, not for decoration.

After the shot list, close with the finishing sections. These describe the whole clip rather than any one beat, so they sit at the bottom where they cannot be read as a step.

  • [ACTING]: the observable behaviour, written as in the performance section above.
  • [LIGHT AND IMAGE]: light direction, contrast, palette, lens feel, and grain.
  • [CAMERA]: the one move you want, or the refusals that keep the frame still.
  • [PRODUCTION SOUND]: the bed, the event sounds in order, and anything banned.
  • [NEGATIVES]: name real failure modes, not generic quality words. Identity drift, wardrobe swaps, extra people, broken eyelines.
Worked example, full structure
[REFERENCE USE]
Image1 = Nadia, lead. Keep the face, the freckles, and the dark curls tied back. Ignore the grey studio backdrop.
Image2 = Theo, second character. Keep the face, the beard, and the wire-rim glasses. Ignore the office behind him.
Image3 = the record shop interior. Keep the wooden browsing bins, the yellow pendant lights, and the tiled floor.
Video1 = motion source only. Copy the slow turn of the head and the reach across the bin. Do not use its actors, its clothing, or its location.
Audio1 = Nadia's speaking voice. Nadia speaks in this voice. Keep her lip sync accurate to the line below.

[IDENTITY / CONTINUITY LOCKS]
Exactly two people are in frame for the whole clip. No extras, no passers-by.
Nadia has the face from Image1. Theo has the face from Image2.
Nadia wears a rust corduroy jacket over a cream tee. Theo wears a charcoal knit sweater.
The hero prop is one sleeved vinyl record. It stays in Nadia's hands until she passes it to Theo.

[SCENE]
The record shop from Image3, late afternoon. Low sun comes through the front window on the right and the pendant lights are already on.

[DIALOGUE]
Nadia: "This is the pressing I told you about."

[SCREEN GEOGRAPHY]
Nadia stands on the left at the browsing bin. Theo stands on the right by the counter. They look at each other when they speak and at the record when it changes hands.

[SHOT LIST]
[0-4 seconds] Medium two-shot, locked off. Nadia turns her head toward Theo and reaches across the bin. End with the record lifted clear of the bin in her right hand.
[4-8 seconds] Slow push in to a medium close-up on Nadia. She says her line and holds the record up. End with the sleeve fully readable in frame.
[8-12 seconds] Hold the medium close-up, then reframe right as she passes the record. End with the record in Theo's hands and Nadia's hands empty.

[ACTING]
Nadia's gaze stays on Theo through the line and drops to the record as she passes it. Her thumbs stay on the sleeve edges. Theo's shoulders drop as he takes it and he exhales once.

[LIGHT AND IMAGE]
Warm low sun from the right, soft fill from the pendants, deep shadows in the bins. Muted rust and amber palette. Shallow depth of field on a 40mm lens, fine grain.

[CAMERA]
One slow push in and one reframe, both on a tripod head. No handheld, no zoom, no whip pan.

[PRODUCTION SOUND]
Quiet shop room tone, the rustle of sleeves in the bin, one soft card-stock slide as the record leaves its slot, and Nadia's line clean over the top. No music.

[NEGATIVES]
No identity drift between Nadia and Theo. No wardrobe changes. No third person entering frame. No broken eyelines. No subtitles or captions.

Read it back before you spend the credits

  • Does every attached asset have exactly one named job, and are the parts you do not want ruled out by name?
  • Does every identity, voice, prop, and edit target have exactly one owner? Two owners means the model picks.
  • Are the timed ranges consecutive, non-overlapping, and long enough for the action inside them?
  • Do the camera, light, and acting sections contradict anything the references already establish?
  • Are the negatives real failure modes rather than words like "bad" or "low quality"?

PixelDojo exposes image references today, so Image1 through Image9 are the tokens you can actually attach. The Video1 and Audio1 contracts are documented here because the model reads them, and because the same discipline of one job per asset is what makes the image-only prompts hold together.

First and Last Frame

Attach a start frame and H3 animates from it. Attach a start and an end frame and it fills the motion between them, which is the most precise control the model offers. Aspect ratio comes from the image in this mode, so the ratio picker is ignored.

Start frame only

The supplied frame is held and animated: he stays at the post and turns his head, the rope moves in the wind.

Start frame and end frame

Same opening frame. He walks out to the left and the clip resolves exactly onto the supplied closing frame, empty dock and coiled rope included.

  • Write the prompt as the motion between the two frames, not as a description of either one. The frames already say what things look like.
  • Use it for transitions where you know both ends: a poster that animates, a product that rotates into a hero angle, a character who exits a set.
  • Frames and reference images are mutually exclusive. Pick one or the other.

Camera Language

H3 reads cinematography vocabulary directly, and it defaults to movement. If you say nothing about the camera, expect a slow drift and continuous reframing. The reliable way to stop that is not to ask for a static shot but to list the moves it should not make.

Nothing said about the camera

Continuous drift. The framing changes throughout and the subject is reframed as she works.

"The frame never moves", with the moves named

One fixed frame for the whole clip. Only the gardener and the spray move. Same seed as the silent version.

Locked-off template
Locked-off static <wide / medium / close> shot.
No push in, no handheld, no zoom, no dolly. The frame never moves.

When you do want a move, name one and describe what changes on screen. "Slowly push in to a close-up of the espresso stream" was executed as written. Bare film-school terms with nothing visible attached are the weakest form of camera direction on every model we have tested, this one included.

Performance and Emotion

Emotion words are a summary of a performance, not a description of one. Write what a camera could see instead: where the eyes go, what the hands do, what the breathing does, what stays still.

"She looks anxious and nervous"

A generic to-camera performance. The violin sits in her lap and the direction resolves into an expression rather than a behaviour.

The same beat written as observable behaviour

Gaze fixed on the floor, fingers gripping the neck of the violin, shoulders held up. The instruction is followed literally and holds for the whole clip.

Performance template
<Subject> <is in the situation>.
<What the hands do>, <where the eyes go>, <what the shoulders or posture do>, and <one breath or pause>.
<Shot size>, <camera behaviour>.

Directing the Sound

Every clip arrives with a stereo track, so sound is something you direct rather than something you receive. Name the sounds in the order they happen, say what runs underneath, and say what should not be there.

No audio direction

A continuous bed. Measured across the clip it sits around −32 dB with little variation.

Sounds named in order, with "no music"

The same picture, but the track is built out of events: quiet floors near −45 dB with transients up to −23 dB where the strikes and the quench land.

Audio template
Audio: <what runs underneath for the whole clip>, <the specific event sounds in the order they happen>, and at <the moment> <the sound that marks it>.
No music.

Levels track the scene, so quiet scenes come back genuinely quiet. Our lakeside and backstage clips measured around −44 dB average, roughly 25 dB below the busy scenes. Expect to ride the gain in post rather than assuming the export is broadcast-ready.

One honest limitation: we verified that every clip carried a real audio track and measured how the levels behaved, but we did not check by ear whether spoken dialogue matched requested words. Treat speech and lip sync as untested here rather than endorsed.

Negative Direction

There is no negative prompt field. Exclusions go in the prompt itself, as plain sentences. We tried twice to show that they do something and could not.

With a negative list: no people, no pigeons, no cars, no text

An empty terrace, as asked.

Same scene, same seed, no negative list

Also an empty terrace. The one difference we could find: this version rendered lettering on the shopfront, which the other did not.

The earlier boat-on-a-lake pair behaved the same way. So the practical advice is narrower than "negatives are unusually effective": use them for the things the model likes to add on its own, which in our testing meant on-screen text and camera movement. Do not count on a negative to remove a subject the scene implies. If a person should not be in the shot, describe an empty place rather than a place with the person banned.

Framing, Length, and Cost

Six aspect ratios, all rendered at 2K: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Vertical is a real 1440 by 2560 rather than a cropped landscape, so it is worth generating social assets natively instead of reframing afterwards.

9:16 at 2K, generated vertical

A product ad written for a phone screen: the camera rises, condensation runs, and nothing is cropped from a wider frame.

SettingRangeNotes
Duration5 to 15 secondsWhole seconds. Cost scales with length.
Resolution2K onlyNo lower tier to save credits on.
Credits3 per second15 credits for a 5-second clip, 45 for the full 15.
Prompt lengthUp to 7,000 charactersA full shot list with sound design fits comfortably.
Aspect ratioSix ratiosIgnored in first-frame mode, where the image sets the shape.

Because every generation is 2K, a rejected clip costs the same as a good one. Test the look and the sound at 5 seconds, then re-run the prompt you like at the length you need.

Checklist Before You Generate

  • Is the camera specified? If not, expect it to drift. If you want it still, list the moves it should not make.
  • Is any on-screen text typed out exactly, with the treatment and position named?
  • Is the performance written as behaviour a camera could see, rather than as an emotion?
  • Does each timed beat carry one primary change and an end state a viewer could point at?
  • Is the last beat one you can afford to lose? It is the one that gets compressed.
  • Is the audio named in order, with anything unwanted ruled out?
  • If references are attached, are the features that must survive named in words as well as shown?
  • Is the length right? You pay 3 credits a second either way, so prove the prompt at 5 seconds first.

Tested on 2026-08-03 against the live model on PixelDojo: 24 generations, run as matched pairs with the same seed and one variable changed. Where a technique made no measurable difference, this guide says so rather than repeating it.