MiniMax H3 AI Generator
Generated on PixelDojo with MiniMax H3. Produced by PixelDojo's generation pipeline.
MiniMax H3 costs 3 credits per second of output at its 2K tier, and a 768p tier exists at 2 credits per second, but only for clips that start from a text prompt with no image. Our 5 second image-to-video test, the shortest duration the model accepts, cost 15 credits and took about 205 seconds on August 26, 2026. We animated a product still of an earbuds case and kept the first result.
The 15 credit clip we timed
Every example below was produced on PixelDojo. Hover to see the prompt.
Slow cinematic camera orbit around the earbuds case, soft studio lighting shifts, the lid opens gently to reveal the earbuds, product commercial style
MiniMax H3
A model that needs you to bring your own audio
Switch models without switching tools. Each one runs in the same PixelDojo studio.
The pricing rule that catches people off guard
Starting from an image locks you into the 2K rate
The provider will not accept a start frame alongside explicit width or height, so an image-to-video job always renders and bills at the 2K tier, 3 credits per second, no matter which resolution you pick in the request. The cheaper 768p tier, at 2 credits per second, is only reachable from a pure text-to-video call with no image.
What the 5 seconds actually showed
Our source photo already had the earbuds case open with both buds visible, so the instruction to have the lid open gently had nothing left to do. The clip instead reads as a slow camera drift with a subtle lighting shift across the case, not a literal opening motion, because the starting frame was already past that point.
The file carries a real audio track
The output is an AAC audio stream muxed into the MP4 alongside the video, matching the model's native synced audio claim. We can confirm the track exists and runs the length of the clip; judging what it sounds like is outside what a still frame or a file probe can tell you.
One endpoint, three modes
Send only a prompt for text-to-video, add image_url for image-to-video, or attach up to 9 reference images plus 3 reference videos plus 3 reference audios for reference-to-video. All three run through the same request, priced the same per second.
Duration is the main cost lever
5 to 15 seconds, integer only. Our 5 second run at 15 credits is the cheapest a 2K clip gets; the same job at 15 seconds would run 45 credits.
Aspect ratio is inherited, not chosen, in image mode
The aspect_ratio field is documented but ignored once an image_url is present. Our square product photo produced a square 1440x1440 output, matching the source rather than any explicit setting.
One product photo, 5 second minimum duration, first result kept. 15 credits, about 205 seconds, native audio track included.
Why Choose Pixel Dojo for MiniMax H3
Professional-quality results with cutting-edge AI technology
3 credits per second at 2K
Flat per second of output duration. A 5 second clip is 15 credits, a 15 second clip is 45.
About 205 seconds, measured
Timed on August 26, 2026 from request to finished file for a 5 second image-to-video clip, returned through an async job poll.
Native synced audio in the same call
No separate audio generation step. The rendered file already has an audio track when the job finishes.
How It Works
Reproduce the 15 credit clip:
Decide if you need an image at all
Adding image_url locks the job into the 2K billing rate. Skip it for a cheaper 768p text-to-video render if a start frame is not essential.
Start at the 5 second minimum
Duration is the main cost lever at 3 credits per second. 5 seconds is the cheapest full clip the model accepts.
Submit and expect a job poll
From the tool page or a POST to /api/v1/models/minimax-h3/run. Ours charged 15 credits and took about 205 seconds, returning through a status check rather than the first response.
Loved by creators on PixelDojo
Real feedback from people using PixelDojo, pulled from our in-product surveys.
exceptional quality and great overall design of platform and interface. very intuative. love the creative freedom.
Qwen image 2 is amazing!!
Creative freedom, range of tools and options.
I love the training feature
the quality is the best
All the tools needed in one dashboard
Explore more AI tools on PixelDojo
AI Tools
Compare & Switch
- Best AI Image Generators
- Best AI Video Generators
- Midjourney Alternatives
- Civitai Alternatives
- Runway Alternatives
- Leonardo Alternatives
- Pika Alternatives
- Luma Alternatives
- Magnific Alternatives
- Veo Alternatives
- Flux Alternatives
- Freepik Alternatives
- Seedance Alternatives
- Seedream Alternatives
- Pixverse Alternatives
- GPT Image Alternatives
- Synthesia Alternatives
- Playground Alternatives
- NightCafe Alternatives
- Canva AI Alternatives
- ElevenLabs Alternatives
- ComfyUI Alternatives
- Fal Alternatives
- Replicate Alternatives
Common Questions
Everything you need to know about MiniMax H3
How much does MiniMax H3 cost?
3 credits per second at the 2K tier, or 2 credits per second at 768p for text-to-video jobs with no starting image. Our 5 second image-to-video run cost 15 credits, billed at the 2K rate because a start frame forces that tier regardless of the resolution setting you send.
How fast is MiniMax H3?
About 205 seconds for a 5 second clip in our August 26, 2026 run. It ran past the API's synchronous window and returned through a job poll, which is normal for video generation.
Why did my image-to-video job cost more than I expected?
Any job that includes a start frame is billed at the 2K rate, 3 credits per second, even if you set resolution to 768p in the request. The provider will not accept explicit dimensions alongside a start frame, so the cheaper tier is only available to text-to-video jobs with no image.
Does MiniMax H3 generate its own audio?
Yes. The rendered file comes back with a synced audio track already in place, no separate text-to-speech or sound generation step required.
How is it different from OmniHuman?
MiniMax H3 generates audio itself as part of the render. OmniHuman needs you to supply an actual audio file as a required input, whether recorded or generated separately, and is built specifically for lip sync on a face rather than open-ended scene animation.
How do I call MiniMax H3 from the API?
POST to /api/v1/models/minimax-h3/run with your API key, a prompt, and optionally an image_url for image-to-video. Duration defaults to 5 seconds, the setting we used.