Skip to main content

Video Reference Images

Generated on PixelDojo. Produced by PixelDojo's generation pipeline.

Cancel anytimeCommercial-use license50+ AI models

Nine video models on Pixel Dojo accept reference images, meaning several pictures that describe a subject rather than one picture the model has to animate. The caps run from three references on VEO 3.1 Standard up to thirty on Seedance 2.5. This is a different input from the start frame that most image-to-video models take, and mixing the two up is the most common reason a request comes back rejected. Both lists are below, with the field name each model expects, read from the catalog on August 26, 2026.

The nine models that take true reference images

Seedance 2.5: up to 30 reference images

The largest reference budget in the catalog by a wide margin, sent as reference_images. It also takes reference videos and reference audio in the same request, and the prompt can address them by name so you can tell the model which reference to use for what.

WAN 3.0: up to 10 reference images

Sent as reference_image_urls, alongside optional reference video and reference audio arrays of up to five each. The prompt syntax lets you point at a specific reference, so one call can carry a face, a set and a voice.

MiniMax H3: up to 9 reference images

Sent as reference_images to guide subject and scene identity. Worth knowing before you build a payload: references cannot be combined with the start frame and last frame fields on this model. You pick one approach per request.

Seedance 2 Reference: up to 9 reference images

A reference-only entry point, sent as reference_images. The request is rejected unless at least one reference image or reference video is present, so there is no accidental text-to-video fallback.

Happy Horse Reference: 1 to 9 reference images

Sent as reference_urls, with at least one required. Clip length runs from 3 to 15 seconds. Version 1.1 bills 3 credits a second at 720p and 4 at 1080p.

Grok R2V: 1 to 7 reference images

Sent as reference_urls, at least one required, clips of 1 to 10 seconds. A tight cap, but enough for a character plus a couple of scene cues.

WAN Reference to Video: up to 5 references

Sent as reference_urls, and the five slots are shared between images and clips, with a maximum of three of those being video. Durations are 5 or 10 seconds. The flash model is 1 credit a second at 720 and 1.5 at 1080; the standard and latest models are 2 and 3.

Kling Reference to Video: up to 4 reference images

Sent as image_urls. This tool has a real mode switch between image references and video references, and the mode changes the rate: image references are 3 credits a second on the standard tier and 7 on pro, while video references are 6 and 8.

VEO 3.1 Standard: up to 3 reference images

Sent as reference_images, for style guidance. The important caveat is that only the Standard tier accepts them. VEO 3.1 Fast and VEO 3.1 Lite do not, so a reference payload has to go to the 8 credits a second tier.

The other kind of image input: a single start frame

Most image-to-video models take exactly one picture and animate it. WAN 2.7, WAN 2.6, WAN 2.2, Vidu Q3, Seedance 1.5, Seedance 2, Grok Video, Flux 3 Video and Gemini Omni Flash all read a single image field. Hailuo 2.3 names its field first_frame_image. Kling Video takes start_image_url and can also take end_image_url to fix the last frame, and PixVerse V6 does the same with image and last_frame_image.

Thirty references on Seedance 2.5, ten on WAN 3.0, nine on MiniMax H3, three on VEO 3.1 Standard. Same idea, very different room to work in.

Why Choose Pixel Dojo for Video Reference Images

Professional-quality results with cutting-edge AI technology

References describe a subject, a start frame describes a shot

Send references when the character or product has to survive into a scene you are describing in words. Send a start frame when you already have the exact opening image and want motion added to it.

The cap tells you what the model is for

A three image cap is style guidance. A nine or ten image cap is identity work. A thirty image cap is closer to a small asset library, which is why the highest caps sit on the models built for multi shot sequences.

The field name is part of the answer

reference_images, reference_image_urls, reference_urls and image_urls all mean the same thing to a person and different things to a server. The field name below is the one that model's schema actually validates.

How It Works

Getting a reference driven clip right on the first attempt:

1

Decide which input you actually need

If you want your exact picture to move, use the start frame field. If you want the subject from your pictures to appear in a scene you are writing, use the reference field. On MiniMax H3 these two are mutually exclusive, so choosing up front avoids a rejected request.

2

Stay inside the model's cap

Three on VEO 3.1 Standard, four on Kling Reference to Video, five on WAN Reference to Video, seven on Grok R2V, nine on Happy Horse Reference, MiniMax H3 and Seedance 2 Reference, ten on WAN 3.0, thirty on Seedance 2.5. Over the cap the schema rejects the call before anything is charged.

3

Name the references in the prompt where the model supports it

WAN 3.0, Seedance 2.5 and Seedance 2 Reference all let the prompt address a specific reference. Saying which image the character comes from is more reliable than hoping the model infers it.

Loved by creators on PixelDojo

Real feedback from people using PixelDojo, pulled from our in-product surveys.

exceptional quality and great overall design of platform and interface. very intuative. love the creative freedom.
Verified PixelDojo creator
Qwen image 2 is amazing!!
Verified PixelDojo creator
Creative freedom, range of tools and options.
Verified PixelDojo creator
I love the training feature
Verified PixelDojo creator
the quality is the best
Verified PixelDojo creator
All the tools needed in one dashboard
Verified PixelDojo creator

Common Questions

Everything you need to know about Video Reference Images

Which AI video model accepts the most reference images?

Seedance 2.5, at up to 30 reference images in a single request. WAN 3.0 is second at 10, and MiniMax H3, Seedance 2 Reference and Happy Horse Reference all cap at 9.

Does VEO 3.1 accept reference images?

The Standard tier does, up to three of them, for style guidance. VEO 3.1 Fast and VEO 3.1 Lite do not accept reference images at all, so a reference driven VEO request has to run on Standard at 8 credits a second.

What is the difference between a reference image and a start frame?

A start frame is the first frame of the clip, so the model animates that exact picture. A reference image describes a subject the model should build into a new scene, and the composition comes from your prompt instead. Models that take references usually accept several; start frame models take one.

Can I use both a reference image and a start frame in the same request?

It depends on the model, and on MiniMax H3 you cannot. Its reference array is documented as mutually exclusive with the start frame and last frame fields. Models like Seedance 2 Reference are reference only by design and reject a request with no reference at all.

Which video models let me fix the last frame as well as the first?

Kling Video takes end_image_url alongside start_image_url, PixVerse V6 takes last_frame_image alongside image, and MiniMax H3 takes end_image_url alongside image_url. That is the way to control where a clip lands, not just where it starts.

Is there an extra charge for sending reference images?

Not on most models. The per second rate is what you pay, and references do not add to it. Kling Reference to Video is the exception worth flagging, because switching from image references to video references changes the rate from 3 credits a second to 6 on the standard tier.

Ready to Create Amazing Video Reference Images Images?

Join thousands of creators using AI to bring their ideas to life