Model grammar profile · Video-01 / H3
MiniMax (Video-01 / H3)
Timed atmospheric direction with multimodal reference roles
Balance rich cinematic atmosphere with precise temporal pacing. Render each meaningful beat in chronological [start-end] sequence blocks, and keep reference roles and native audio as typed multimodal inputs.
Recommended sentence order
01subject→
02context→
03style→
04action→
05cinematography→
06dialogue→
07audio
Negative prompt policy
Use exclusions the way this surface expects.
No formal MiniMax negative-prompt grammar is documented in the public H3 guide. Keep exclusions as concise scene constraints and avoid importing diffusion-style weights or comma blacklists.
Typed parameters
Duration4 · 6 · 8 · 10 · 15
Aspect ratio16:9 · 9:16 · 1:1
Generation modeText to video · First/last frame · Reference generation
Reference rolesUser supplied
Native audio / dialogueOn / off
Reliability constraints
- MiniMax H3 accepts text, image, video, and audio together; attach assets as typed reference roles rather than describing hidden source files in prose.
- Place camera control immediately after the key visual beat; official examples use direct controls such as [pan], [zoom], or [static].
- Use an H3 Context-IR pass only as an explicit enhancement workflow because it is asynchronous and returns an enhanced prompt rather than a video.
Source material
MiniMax video generation guide MiniMax image-to-video API