Text to video
Direct a full scene from a text brief, with prompts up to 7,000 characters.
Independent launch guide · Updated July 31, 2026
Verified specifications, launch analysis, prompt craft, and availability updates for MiniMax’s new general-purpose multimodal video model.
Verified July 31, 2026. Specifications on this site are checked against MiniMax's release notes and technical documentation. Unknowns stay labeled as unknowns.
2K
maximum output
4–15s
integer duration
12
mixed references
7,000
prompt characters
One model, more context
MiniMax describes H3 as a general-purpose model that understands text, image, video, and audio inputs in a unified generation workflow.
Direct a full scene from a text brief, with prompts up to 7,000 characters.
Use first and last frames or image references to anchor composition and identity.
Combine up to 12 total reference files, subject to per-media limits.
Provide audio references alongside image and video inputs for generation direction.
Demo desk
We do not download or re-upload creator work. Until official platform embeds are available, the demo desk links directly to MiniMax’s documented examples and keeps analysis separate from launch claims.
Official documentation
View the original examples in context, with MiniMax's input limits and API parameters alongside them.
Every future embed will retain creator attribution, original link, platform controls, and an editorial note.
Confirm the model name and release status at the source.
Access status
The fastest path to trust is saying exactly what works today.
API documented; this site is not the generator
MiniMax documents MiniMax-H3 on its video generation API. MiniMaxH3.xyz currently provides independent guidance and a local prompt builder—it does not upload media or send generation jobs.
Build with clarity
Compose scene, camera, lighting, motion, and audio direction in one structured brief.
Open prompt builder