MiniMax H3 is MiniMax's general-purpose multimodal AI video generator for text-to-video, image-to-video (first / first-last frame), and Omni Reference workflows, plus V2V motion transfer and video editing. Export 768P (default) or 2K clips at 4–15 second durations. Text-to-video and multimodal modes support adaptive, 16:9, 9:16, and more; first-last frame image-to-video does not offer ratio selection. Built-in native stereo audio makes it a strong fit for ad videos, brand films, ecommerce motion assets, product design, UI/UX demos, and game previs.
Combine up to 9 images, 3 videos, and 3 audio files (12 total); audio refs require an image or video partner to reduce drift and off-brief outputs
Built-in stereo soundtracks with 768P default and optional 2K for marketing and brand clips that need near-final polish
H3 is MiniMax's newer line with full Omni Reference, native stereo, 2K, editing, and V2V motion transfer for complex creative jobs
Text, image, multimodal, and V2V in one picker—ads, branding, ecommerce, product design, UI/UX, and game previs without model hopping
Every scenario requires a non-empty prompt (up to 7,000 characters). For text-to-video or multimodal mode, upload up to 9 images (≤30MB, 256–5760px, JPG/PNG/WEBP/HEIC), 3 videos (≤50MB, 2–15s, MP4/MOV), and 3 audio files (≤15MB, 2–15s, WAV/MP3). For image-to-video or first-last frame, upload one or two stills (≤30MB, 256–5760px).
Explore KreadoAI text-to-video, image-to-video, and multi-model AI video generation to produce iterable 2K AI video assets for marketing and ecommerce teams
Pick MiniMax H3 in KreadoAI—text-to-video, first-last frame image-to-video, and Omni Reference for ad and ecommerce AI video production