Image 1
= face
Image 2
= product
Video 1
= camera
Audio 1
= voice
Label each reference
Tag every file with a role. H3 fuses image, video, and audio — vague labels waste generations.
Hailuo3 AI runs MiniMax-H3 (Hailuo 03 / Hailuo 3.0): text-to-video, first- and last-frame control, and multimodal references. Output is native 2K with stereo audio, 4–15 seconds per clip — generate and iterate in the browser at hailuo3-ai.com.
Examples of H3-class output. Recreate a similar brief in the workspace with your own prompt and references.
H3 is MiniMax’s general-purpose multimodal video model. It reads text, images, video, and audio together, generates picture with native stereo sound, and supports instruction-based editing — so you can brief once and refine instead of restarting from zero.
API resolution is 2K by default. H3 regenerates detail in-context rather than bolting on a separate upscaler, which helps small text and fine product detail survive the master.
Try it nowOfficial H3 generations include stereo audio with the picture — dialogue, effects, and atmosphere in one audiovisual pass, not a silent video you dub later.
Try it nowChoose any whole-second length from 4 through 15. Long enough for a multi-beat hook; short enough to iterate without waiting on a full short film.
Try it nowIn reference mode: up to 9 images, 3 videos (2–15s each, 15s total), and 3 audio clips (audio must pair with an image or video). Mix subject, motion, style, and voice in one brief.
Try it nowH3 supports precise multimodal editing: change characters, objects, scenes, sound, or pacing with language instead of only regenerating from a blank prompt.
Try it nowOfficial demos cover product UI motion, game-style camera, and stylized animation — useful for previz, PV mood, and interface storytelling.
Try it nowFrom account to export

Create an account at hailuo3-ai.com. The product is focused on MiniMax-H3 — you are not picking from a long multi-model menu.

Text-to-video needs a prompt and an explicit aspect ratio. Image-to-video uses first and/or last frames. Reference mode adds images, video, and optional audio under the official caps.

Pick 4–15 seconds at 2K, generate, preview with audio, then download the result or describe an edit instruction for another pass.

Generate sound-on variants for the same offer. Use reference images for product and talent so color and identity stay closer across takes.
Try it Now
Combine mood, character, product, and logo references in one multimodal brief — the pattern MiniMax highlights for premium brand films.
Try it Now
Vertical trailers and character-led beats benefit from face/voice references and multi-shot modeling inside a single 4–15s generation.
Try it Now
Native 2K masters for pack shots, feature demos, and listing media where still photos under-sell the product.
Try it NowThree things to set before you hit generate
Image 1
= face
Image 2
= product
Video 1
= camera
Audio 1
= voice
Tag every file with a role. H3 fuses image, video, and audio — vague labels waste generations.
Audio needs an image or video pair — never alone.
Stay under the official caps. Audio cannot go alone — always pair it with image or video.
Aspect ratio
Pick 4–15 seconds. Text-to-video needs a ratio; image-to-video follows the frame; reference mode can stay adaptive.
If you already want H3, skip the multi-model maze
Need sound-on 9:16 and 16:9 variants of the same offer without configuring an API console.
Want multimodal references (look, talent, product, logo) in one brief before a live shoot.
Need 2K product motion that still holds on listing pages after export.
Prefer instruction edits on a near-miss take over spending another full generation from scratch.
Use face and voice references so the same character can hold across short vertical beats.
Prototype UI motion or game-style camera with stereo audio already in the draft.
Need sound-on 9:16 and 16:9 variants of the same offer without configuring an API console.
Want multimodal references (look, talent, product, logo) in one brief before a live shoot.
Need 2K product motion that still holds on listing pages after export.
Prefer instruction edits on a near-miss take over spending another full generation from scratch.
Use face and voice references so the same character can hold across short vertical beats.
Prototype UI motion or game-style camera with stereo audio already in the draft.
Need sound-on 9:16 and 16:9 variants of the same offer without configuring an API console.
Want multimodal references (look, talent, product, logo) in one brief before a live shoot.
Need 2K product motion that still holds on listing pages after export.
Prefer instruction edits on a near-miss take over spending another full generation from scratch.
Use face and voice references so the same character can hold across short vertical beats.
Prototype UI motion or game-style camera with stereo audio already in the draft.
Model facts and the Hailuo3 AI product
Hailuo3 AI (hailuo3-ai.com) is a web product focused on MiniMax-H3 — the model also called Hailuo 03 or Hailuo 3.0. It is not the MiniMax official console; it packages H3 generation for creators who want a single-model workspace.
Every job needs a text prompt (up to 7,000 characters). Reference mode allows ≤9 images, ≤3 videos, ≤3 audio clips, ≤12 files total. Audio references must include an image or video. Per-file caps: image 30 MB; video 50 MB; audio 15 MB. Video reference clips are 2–15s each with total ≤15s.
Yes. MiniMax describes H3 as generating video with native stereo audio. Preview with sound on before you export from Hailuo3 AI.
MiniMax-H3 is MiniMax’s general-purpose multimodal video model (API model name MiniMax-H3). It accepts text, image, video, and audio context, outputs native 2K video with stereo audio, supports 4–15 second integer durations, and offers text-to-video, first/last-frame image-to-video, reference generation, and instruction-based editing.
Output resolution is 2K. Duration is any integer from 4 to 15 seconds. Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive where the mode allows. Text-to-video requires a fixed ratio (not adaptive). First/last-frame image-to-video follows the input image framing.
Hailuo 2.3 (MiniMax-Hailuo-2.3) is a prior video line with shorter clips and up to 1080p on documented API settings (for example 6s/10s at 768P, 6s at 1080P). MiniMax-H3 targets native 2K, 4–15s, unified multimodal references (image/video/audio), native stereo audio, multi-shot modeling, and instruction-based editing in one general-purpose model.
Create a free account, open the MiniMax-H3 workspace when generation is available on your plan, and upgrade if you need higher volume.
Create accountCredit plans for MiniMax H3 (Hailuo 03). See how many 5s / 10s clips each tier covers.
Credits refill every billing cycle. Video estimates use MiniMax H3 at 5s / 10s · 2K.
Start experimenting with the magic of AI video.
Unlock deeper control and create at scale.
Create at full speed with uninterrupted flow.
Imagine without limits. Make your wonder.