MiniMax H3

MiniMax H3

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.

About

MiniMax H3

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.

MiniMax H3
How it works

Create with MiniMax H3

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.

01

Describe the shot

Write the subject, action, camera movement, lighting, and sound you want in one clear prompt.

02

Add references

Use an image, video, or audio reference when the subject, motion, style, or rhythm needs tighter control.

03

Generate and refine

Choose the duration, aspect ratio, and resolution, then iterate on the prompt until the shot is ready.

Duration

4 / 6 / 8 / 10 / 12 / 15

Resolution

768p / 2k

Aspect Ratio

adaptive / 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16

Reference Image

Max 9 images for video generation

MiniMax H3 — Settings

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.

Text

Suitable Scenes: quick creative drafts, product motion, social clips, and projects using multiple reference media.

image

Prompt tips: State the main motion clearly and explain whether each reference controls the subject, movement, or rhythm.

and reference-to-video

Social clips,Product animation,Creative drafts,Reference video tests

Examples & use cases

What you can make with MiniMax H3

Explore practical ideas for camera direction, reference control, character consistency, native audio, art direction, and precise prompt timing.

01

Camera & motion

Camera moves you can direct

Describe a tracking, orbit, tilt, or push-in shot and keep the movement tied to the subject, background, and pacing.

View example prompt

A character walks across a beach in golden light with natural camera movement and realistic fabric motion.

02

First & last frame

Two references become one shot

Use a start frame and an optional end frame when the opening and closing composition need to stay intentional across the transition.

View example prompt

Let the product from the reference image slowly rotate on a black table.

03

Character consistency

The same character, shot after shot

Keep the character identity, clothing, proportions, and visual traits stable while the location, lighting, or camera changes.

View example prompt

Create a short social clip with movement timed to the reference audio rhythm.

04

Native audio

Sound generated with the picture

Describe dialogue, ambience, foley, music, or rhythm together with the shot so the audio direction follows the visual action.

View example prompt

A short cinematic shot with synchronized ambience, foley, and a clear audio rhythm that matches the movement on screen.

05

Art direction

One look held across every shot

Give the model a visual language—palette, lighting, linework, or material—and carry it consistently through the whole clip.

View example prompt

A sequence with a consistent palette, lighting language, texture, and art direction from the first frame to the last.

06

Prompt adherence

It hits the beats in the order you write them

Put the subject, action, camera, timing, and constraints in a clear order; precise prompts make complex shots easier to control.

View example prompt

A precise shot brief with ordered actions, camera direction, timing, lighting, and output constraints.

FAQ

MiniMax H3 FAQ

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.

Explore More AI Video Models

MiniMax H3

MiniMax H3 is a multimodal video model for text, image, and reference-to-video generation with 768p/2K output, flexible duration, and reference audio.