MiniMax H3
Open-weights video model from MiniMax generating native 2K clips up to 15 seconds with synchronized audio and multi-modal reference conditioning.
Updated 2026-08-01
Overview
MiniMax H3 is the Chinese lab's next-generation video model, released on July 31, 2026. It generates clips at native 2K up to roughly 15 seconds, and the headline capability is that it takes text, image, video and audio as inputs to the same prompt rather than making you pick a mode. You can hand it a reference image for a character, a reference clip for motion or framing, and a text instruction for what should happen — the model resolves all of them together. MiniMax calls this omni-reference; in practice it's the feature that separates H3 from the text-to-video-plus-optional-first-frame pattern most competitors still use.
The second thing that matters is audio. H3 produces sound synchronized to the generated footage in the same pass, which puts it in the small group of models — Veo 3 being the obvious comparison — that don't leave you compositing a soundtrack afterwards. Native 2K also means you aren't upscaling 720p output before it's usable anywhere near a real timeline. It's already available through third-party inference hosts including fal.ai, so you can hit it via API without going through MiniMax's own platform.
The audience here is developers and studios building video generation into a product, not casual users looking for a web app to play with — this is a model release first. Open weights are the strategic differentiator: if you can run it yourself, you can fine-tune on your own footage, keep client material off a third-party server, and avoid per-generation API costs at volume. That's a meaningfully different proposition from Veo 3 or Kling, both of which are closed. The tradeoff is that a 2K video model with audio is not something you casually self-host, and MiniMax's tooling and documentation remain thinner than what US labs ship.
Key features
Omni-reference inputs
Accepts text, image, video and audio references in a single prompt and reconciles them, so you can lock a character's appearance from a still while borrowing camera motion from a clip. Most rival models handle one conditioning signal at a time.
Native 2K generation
Outputs at 2K resolution directly rather than generating at 720p or 1080p and upscaling. Fewer artifacts survive into the final frame, and the footage drops into a real edit without an intermediate enhancement step.
Synchronized audio
Generates a matched audio track alongside the video in the same pass instead of leaving sound design as a separate job. Puts H3 in the same category as Veo 3 and ahead of most open video models, which are silent.
Open weights
The model is released with downloadable weights, so teams can self-host, fine-tune on proprietary footage, and keep sensitive material off external APIs — a route that closed models like Veo 3, Kling and Runway simply don't offer.
Pricing
Free tier: Yes — a limited free tier is available for evaluation. Check minimax.io for current generation limits.
| Plan | Price | What's included |
|---|---|---|
| Free tier | $0 | Limited trial generations through MiniMax's platform to evaluate output quality before committing to API spend. |
| API (token-based) | Usage-based | Pay-per-generation pricing through MiniMax's platform and third-party hosts such as fal.ai. Rates vary by resolution and clip length — check the official pricing page for current numbers. |
| Self-hosted | Your own compute | Open weights can be downloaded and run on your own GPUs with no per-generation fee. Hardware requirements for 2K video with audio are substantial. |
Limited trial generations through MiniMax's platform to evaluate output quality before committing to API spend.
Pay-per-generation pricing through MiniMax's platform and third-party hosts such as fal.ai. Rates vary by resolution and clip length — check the official pricing page for current numbers.
Open weights can be downloaded and run on your own GPUs with no per-generation fee. Hardware requirements for 2K video with audio are substantial.
Pros & cons
Pros
- ✓Open weights make self-hosting and fine-tuning possible — rare for a video model at this capability level
- ✓Combines text, image, video and audio references in one prompt instead of forcing a single conditioning mode
- ✓Native 2K output skips the upscaling step most generators still require
- ✓Generates synchronized audio in the same pass, not as a separate job
- ✓Already served by third-party inference hosts like fal.ai, so you can test it without onboarding to MiniMax's platform
Cons
- ×The roughly 15-second ceiling means anything longer has to be stitched from multiple generations, with the usual continuity problems
- ×Self-hosting a 2K video model with audio needs serious GPU capacity — the open weights are practically out of reach for individuals
- ×Token-based API pricing makes per-project cost hard to forecast compared to a flat monthly plan
- ×Documentation and developer tooling are thinner than Google's or Runway's, and Chinese-lab origin raises compliance questions for some enterprise buyers
How it compares
| Tool | Best for | Pricing | Score |
|---|---|---|---|
| MiniMax H3 | — | Free tier + token-based API pricing (check site for current rates) | 8.2/10 |
| Runway vs Runway → | Filmmakers, advertisers, and content creators who need cinematic AI video with realistic motion and fine-grained creative control. | Freemium | 9.3/10 |
| Veo 3 vs Veo 3 → | Filmmakers, content creators, and marketing teams who need production-quality cinematic video with synced audio, without a production budget. | Free via Gemini + Vertex AI pay-per-use | 9.1/10 |
| Seedance 2.0 vs Seedance 2.0 → | Video creators and developers who need short AI-generated clips with synchronized audio, lip-sync, and cinematic camera control. | Free daily credits (Dreamina) + paid ~$15-$70/mo; API from ~$0.08/s | 9/10 |
Compare head-to-head
Related reading
DeepSeek-V4-Flash API: Setup, Model ID, Failure Modes
DeepSeek opened the V4-Flash API public beta on July 31, 2026. The shortest path to a working agent call, and the beta failures to plan around.
Anthropic Discloses Three Claude Eval Escapes
Anthropic says Claude models escaped sandboxed cyber evals and reached three organizations' live systems. Where each frontier lab's containment stands.
GPT-5.6 Price Cuts Land Only on the Cheap Tiers
OpenAI cut GPT-5.6 Luna 80% and Terra 20% on July 30. Sol's list price held, and that is the tier long agent runs actually bill against.
Ready to try MiniMax H3?
Head to the official site to start with MiniMax H3 — pricing and plans are listed above.
Visit MiniMax H3

