Early access

MiniMax H3 on Overchat AI

One omni-modal model that reads text, images, video, and audio as a single context — then returns a 4 to 15 second clip at 2K with stereo sound already inside it. Join the list and we'll let you know the moment it's ready.

We'll email you when MiniMax H3 goes live. No spam.

Introducing MiniMax H3

Up to 12 Reference Files

Provide MiniMax with image, video and audio files simultaneously, and incorporate your own elements seamlessly into AI videos. You can provide up to nine reference images, three video clips and three audio clips in a single request — a total of twelve files. Describe how you want MiniMax H3 to use these files in plain language.

15 Seconds of AI Video at 2K

Generate AI-powered video clips in a cinematic style, measuring 4–15 seconds in length, with a resolution of 1440 pixels on the short edge and 24 frames per second. Generate videos in 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 aspect ratios, or use adaptive framing to ensure they are optimised for any device. Alternatively, switch to 768p mode to iterate quickly at a lower cost than competing mainstream AI video generation models.

Native Stereo Audio

The MiniMax H3 supports built-in audio generation, enabling the addition of a stereo track containing dialogue, foley and environmental effects that are already perfectly timed to camera cuts and on-screen action. Simply describe the sounds or what the characters say in your prompt.

Advanced and Reliable Video Editing

In a series of blind tests by the AI community, MiniMas H3 was ranked as the number one favourite model for video editing when it was released. This model offers reliable prompt following, allowing you to edit videos naturally through natural language commands — add sounds and elements, change colours, modify scenes and camera angles; the possibilities are limitless.

Introducing Overchat AI

Overchat AI brings you the power of the world’s top AI models: GPT, Claude, Gemini, Mistral, and more.

Genereate videos with Gemini Veo 3 text-to-video tool on Overchat AI

Use Cases

What can you create with MiniMax H3? Get inspired with these ideas:

📱

Create Short-Form Videos

Generate up to 15 seconds of 2K content for TikToks, Reels and Shorts from a single prompt. Select the 9:16 aspect ratio to make your video ready to post to social media immediately.

🎬

Make AI Films

Describe the scene, mood, action and camera work, then watch as the AI brings your imagination to life from just a text prompt.

🎞️

Edit Videos

Upload a video, describe the changes you want to make, and MiniMax H3 will edit your video just as if you were using Final Cut — except you describe the changes in plain English.

👥

Add Characters to Videos

Import your reference images and videos, then add a video file. MiniMax H3 will then add the character to the target video using either a photo or another video.

🛍️

Add Products to Videos

Creating product shots has never been easier. Simply provide MiniMax H3 with a photo, multiple photos, or even a video of your product, along with a description of the scene. You can then easily create hero shots or short commercials using AI.

🌟

Make Images Come Alive

Use MiniMax H3 to turn your favourite photographs into videos and bring your memories to life. You can also animate your favourite memes, turn sketches and drawings into videos, and much more besides.

How it Works

Create with MiniMax H3 in 3 simple steps

✍️

Describe What You Want

Write your prompt and attach any reference images, clips or audio for the MiniMax H3 to use.

01
🤖

MiniMax H3 Generates It

The model produces your clip and audio. Generation usually takes 30 seconds to a couple of minutes.

02
📥

Download and use

Get your result ready to share, post, or integrate into your projects.

03

FAQ

What is MiniMax H3?

arrow

The MiniMax H3 (also known as the Hailuo 3.0), an omni-modal video model developed by MiniMax, was released on 31 July 2026. It can use text, images, videos and audio references to generate 4–15 second clips at 2K resolution with native audio. It can also use up to 12 reference files from different modalities. This AI video generation model excels in video editing.

What's unique about MiniMax H3 compared to other AI video models?

arrow

In MiniMax H3, all inputs — whether text, image, video or audio — form one shared context that the model can process internally. In most models, these inputs are processed in separate pipelines, one for each media type. This architecture allows the model to represent the reference more accurately in the final generation.

How long can MiniMax H3 videos be?

arrow

The MiniMax H3 can generate videos of 4 or 15 seconds in length at 2K resolution (1440 pixels on the short edge) at 24 frames per second. Aspect ratios include 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 and an adaptive option.

Does MiniMax H3 generate audio?

arrow

Yes, the MiniMax H3 can generate stereo audio clips (with two separate audio channels, left and right) to create immersive environmental soundscapes, natural-sounding dialogue and punchy audio effects. To create audio with MiniMax H3, simply enter a description in the text prompt. If you want to include dialogue, put what the characters should say in double quotation marks, "like this", and use the prompt to describe which piece of dialogue belongs to which character.

Is MiniMax H3 open source?

arrow

MiniMax has announced that the H3 is an open-weight model, meaning that the weights are available under a planned MiniMax Community Licence. This licence allows organisations with revenues of up to $20 million to use the weights commercially, provided they attribute them to MiniMax. While this is not open source, as the learning pipeline and other proprietary aspects of creating the MiniMax H3 will remain closed source, it does mean you can download the weights, fine-tune them and use the model in your own projects, even commercially, as long as your revenue is under 20 million US dollars.

How many reference files can MiniMax H3 take?

arrow

The MiniMax H3 can take an impressive array of nine reference images and three reference video clips, each between two and 15 seconds long (but no more than 15 seconds in total). In total, the model can take twelve files per generation. Please note, however, that you cannot attach audio references without also adding image or video references.