OmniHuman AI Generator
Generated on PixelDojo with OmniHuman. Produced by PixelDojo's generation pipeline.
OmniHuman charges 3 credits per second of the audio clip you give it, up to a 35 second limit, and takes an image plus that audio as its two required inputs. Our test cost 24 credits for a clip billed at 8 seconds and took about 510 seconds to render on August 26, 2026, which is slow next to most image and video models in this catalog. We fed it a documentary style portrait and a short narration line we generated separately, and it turned both into a synced talking clip.
Documentary portrait, animated to speak the line
Every example below was produced on PixelDojo. Hover to see the prompt.
Talking head narration, natural delivery, audio: "Welcome back. Today we are looking at how storm systems form over warm ocean water, and why the spin starts."
OmniHuman
A video model that generates its own synced audio
Switch models without switching tools. Each one runs in the same PixelDojo studio.
What it took to get one talking clip
Two required inputs, not one
OmniHuman needs an image and an audio file, both as URLs. There is no text-only mode. If you only have a script, you generate the audio first, then feed the result in as the second input.
Our two-step setup
We generated a short narration with text-to-speech first: 'Welcome back. Today we are looking at how storm systems form over warm ocean water, and why the spin starts.' That came out to 7.5 seconds of audio for 0.5 credit, the platform's minimum text-to-speech charge. We then fed that file plus a fisherman portrait into OmniHuman.
Price is billed on your audio estimate, not the render
The credit charge comes from the audioDuration you send with the request, used purely for the cost calculation. We estimated 8 seconds and were billed 24 credits. The rendered clip actually came out at 7.56 seconds, matching the real audio almost exactly, but the estimate is what determines the price, so a rough guess can over or undercharge slightly.
It reframed toward the face on its own
Our source photo was a hands-focused documentary shot of a man mending a net, looking down and away from camera. Partway into the clip, the framing shifted upward to bring his face forward and level with camera, mouth moving in time with the narration. OmniHuman does not require a face-forward portrait to start from; it works to get the face into frame regardless.
fast_mode trades quality for speed
We ran with fast_mode on, and the render still took about 510 seconds. This model is simply slower than most of the catalog regardless of that setting, because it is doing full lip sync animation rather than a generic camera move.
35 second ceiling
Audio duration must stay under 35 seconds, which caps a single OmniHuman call at 105 credits. Longer scripts need to be split into separate clips.
One portrait, one narration clip, first result kept. 24 credits for OmniHuman, about 510 seconds, plus a 0.5 credit text-to-speech step to make the audio in the first place.
How It Works
Reproduce the talking head clip:
Generate or record the audio first
OmniHuman has no text mode. We used text-to-speech for a short narration line, which cost 0.5 credit for 7.5 seconds of audio.
Pick any portrait, even a non-ideal one
Our source was a documentary hands shot, not a face-forward headshot, and OmniHuman still brought the face into frame to deliver the line.
Estimate the audio length honestly when you submit
The audioDuration field you send sets the price. From the tool page or a POST to /api/v1/models/omnihuman/run. Ours was billed 24 credits for an 8 second estimate and took about 510 seconds to render.
Loved by creators on PixelDojo
Real feedback from people using PixelDojo, pulled from our in-product surveys.
The best tools for IA on the web !
The guy that operates the website is constantly updating it
I’ve been using PixelDojo for around 8 months now and it’s always worked wonderfully while constantly adding features.
The site is easy to navigate and use.
All the tools, readily available and easy to understand
You guys, the LORA training is the BEST in the space right now, great job with that!
Explore more AI tools on PixelDojo
AI Tools
Compare & Switch
- Best AI Image Generators
- Best AI Video Generators
- Midjourney Alternatives
- Civitai Alternatives
- Runway Alternatives
- Leonardo Alternatives
- Pika Alternatives
- Luma Alternatives
- Magnific Alternatives
- Veo Alternatives
- Flux Alternatives
- Freepik Alternatives
- Seedance Alternatives
- Seedream Alternatives
- Pixverse Alternatives
- GPT Image Alternatives
- Synthesia Alternatives
- Playground Alternatives
- NightCafe Alternatives
- Canva AI Alternatives
- ElevenLabs Alternatives
- ComfyUI Alternatives
- Fal Alternatives
- Replicate Alternatives
Common Questions
Everything you need to know about OmniHuman
How much does OmniHuman cost?
3 credits per second of the audio clip you send, up to a 35 second cap, so the maximum a single call can cost is 105 credits. The charge is based on the audioDuration value you submit with the request, not a measurement of the actual file. We estimated 8 seconds and were billed 24 credits for a clip that rendered at 7.56 seconds.
How fast is OmniHuman?
About 510 seconds in our August 26, 2026 test, for an 8 second billed clip with fast_mode on. That is noticeably slower than most models in the catalog, because it is running full lip sync animation rather than a general camera move.
Do I need to generate the audio separately?
Yes. OmniHuman takes an audio file URL as a required input; it does not accept a text script directly. We generated ours with a separate text-to-speech call, which cost 0.5 credit for our 20 word line, then passed the resulting file into OmniHuman.
Does the source photo need to be a face-forward portrait?
No. Ours was a hands-focused documentary shot with the subject looking down and away from camera. OmniHuman brought the face into frame and toward camera on its own partway through the clip to deliver the line.
How is OmniHuman different from MiniMax H3?
MiniMax H3 can generate its own synced audio directly from a text description in one call, at 3 credits per second of video. OmniHuman needs you to supply the actual audio file yourself, whether recorded or generated separately, and it is built specifically for lip sync animation on an existing face rather than open-ended scene generation.
How do I call OmniHuman from the API?
POST to /api/v1/models/omnihuman/run with your API key, an image URL and an audio URL. Add audioDuration with your best estimate of the audio's length in seconds, since that value sets the credit charge.