One Context for Text, Image, Video, and Audio
MiniMax H3 reads text, images, video and audio as a single context instead of one modality at a time. You can hand it up to nine reference images, three video clips and three audio clips in the same request — twelve files in all — and describe in plain language how they relate to each other: this face, that camera move, this room's sound. MiniMax released it on 31 July 2026; in the Hailuo AI app the same model appears as Hailuo 3.0.
Up to 15 Seconds at 2K
Every generation runs 4 to 15 seconds at 2K — 1440 pixels on the short edge, 24 frames per second — in 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 or adaptive framing. A 768p mode is there for cheap iteration. At 2K the per-second price sits under a third of what mainstream video models charge, and 768p costs less than half of their 720p.
Native Stereo Audio, Not a Separate Pass
The clip arrives with its own stereo track: score, dialogue, foley and room tone, timed to the cut. There is no separate voice, effects and music stage to line up afterwards, because the model never treated them as separate things — the sound and the picture come out of the same generation.
Ranked First for Video Editing
On the Artificial Analysis leaderboards H3 sits first in Video Editing, second in Text-to-Video and third in Image-to-Video, all measured with native audio included. Editing is where it leads: give it existing footage plus a description of the change and it reworks that clip rather than regenerating the shot from nothing.
Overchat AI brings you the power of the world’s top AI models: ChatGPT, Claude, Gemini, Mistral, and more.

What can you create with MiniMax H3? Get inspired with these ideas:
Short-Form Video
Generate TikToks, Reels and Shorts with the sound already inside the clip — up to 15 seconds at 2K from a single prompt, framed 9:16 so nothing needs cropping afterwards.
AI Films
Build scenes that carry motion and atmosphere, with score, dialogue and foley timed to the cut by the same model that generated the picture.
Video Editing
Hand H3 a clip and describe the change. Editing is its strongest ranking — first on Artificial Analysis — because it reworks the footage you gave it instead of regenerating the shot from nothing.
Reference-Driven Casting
Pin a face, a location and a camera move with up to nine reference images, then let H3 hold them steady across the whole clip.
Product Marketing
Turn a product shot into a moving clip with sound, or build a campaign cut from one description plus a handful of reference frames.
Image to Video
Start from a still and let H3 animate it, or give it up to three reference clips and three audio clips and reshape those instead.
Create with MiniMax H3 in 3 simple steps
Describe What You Want
Write your prompt and attach any reference images, clips or audio you want H3 to follow.
MiniMax H3 Generates It
The model produces your clip and its stereo audio together, in one pass through the same network.
Download and use
Get your result ready to share, post, or integrate into your projects.
What is MiniMax H3?
MiniMax H3 is an omni-modal video model released by MiniMax on 31 July 2026. It reads text, images, video and audio as one context and generates 4 to 15 second clips at 2K with native stereo sound. In the Hailuo AI app the same model appears as Hailuo 3.0, and the API labels it Hailuo-03.
What makes MiniMax H3 different from other AI video models?
Two things. Every input — text, image, video, audio — is part of one context rather than routed through its own pipeline, so you can describe reference and editing relationships in plain language instead of picking a mode. And the audio is produced with the picture, so there is no separate sound pass to line up afterwards.
How long can MiniMax H3 videos be?
Between 4 and 15 seconds, in whole seconds, at 2K — 1440 pixels on the short edge, 24 frames per second. Aspect ratios cover 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, plus an adaptive option, and a 768p mode is available for cheaper iteration.
Does MiniMax H3 generate audio?
Yes, and not as a second step. Every clip comes back with a stereo track — score, dialogue, foley and room tone — generated together with the video and timed to the cut, with no split between voice, effects and music.
Is MiniMax H3 open source?
MiniMax announced H3 as open-weight and said the weights would follow in the coming days under a planned MiniMax Community License, which is set to allow commercial use for organisations under $20M in revenue with attribution. As of early August 2026 the weights had not shipped yet.
How many reference files can MiniMax H3 take?
Up to nine reference images, three reference video clips of 2 to 15 seconds each (15 seconds total) and three reference audio clips, capped at twelve files per request. Audio references cannot be submitted on their own.