ai.image.generator.loading
Text to video
01Generate a 5–15 second clip from a prompt with 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 framing.
Generate a 5–15 second clip from a prompt with 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 framing.
Create 5–15 second 480p/768p videos from text, first and last frames, or multimodal reference media. MiniMax H3 Max supports a dedicated reference-to-video workflow with images, video and audio.

MiniMax H3 Max is a video generation model in the H3 family for text-to-video, image-to-video with a first frame, a last frame, or both, and multimodal reference-to-video. The current EvoLink routes support 480p or 768p MP4 output from 5 to 15 seconds. Reference mode accepts images, videos and audio, making H3 Max useful when a shot needs stronger identity, motion or timing guidance than a text prompt alone.

Choose text, keyframes, or multimodal references based on how much control your shot needs.
ai.image.generator.loading
Generate a 5–15 second clip from a prompt with 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 framing.
Generate a 5–15 second clip from a prompt with 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 framing.
ai.image.generator.loading
Animate from a first frame, a last frame, or both when you need a defined visual start and finish.
Animate from a first frame, a last frame, or both when you need a defined visual start and finish.
ai.image.generator.loading
Use up to 9 images, 3 videos and 3 audio clips, with no more than 12 reference assets in one request.
Use up to 9 images, 3 videos and 3 audio clips, with no more than 12 reference assets in one request.
ai.image.generator.loading
Use 480p for lower-cost exploration or 768p when you need a cleaner review-ready result.
Use 480p for lower-cost exploration or 768p when you need a cleaner review-ready result.
MiniMax H3 Max FAQ