MiniMax H3 logo
Video Free tier + token-based API pricing (check site for current rates)

MiniMax H3

Open-weights video model from MiniMax generating native 2K clips up to 15 seconds with synchronized audio and multi-modal reference conditioning.

Updated 2026-08-01

8.2
AI Score / 10
Visit MiniMax H3

Overview

MiniMax H3 is the Chinese lab's next-generation video model, released on July 31, 2026. It generates clips at native 2K up to roughly 15 seconds, and the headline capability is that it takes text, image, video and audio as inputs to the same prompt rather than making you pick a mode. You can hand it a reference image for a character, a reference clip for motion or framing, and a text instruction for what should happen — the model resolves all of them together. MiniMax calls this omni-reference; in practice it's the feature that separates H3 from the text-to-video-plus-optional-first-frame pattern most competitors still use.

The second thing that matters is audio. H3 produces sound synchronized to the generated footage in the same pass, which puts it in the small group of models — Veo 3 being the obvious comparison — that don't leave you compositing a soundtrack afterwards. Native 2K also means you aren't upscaling 720p output before it's usable anywhere near a real timeline. It's already available through third-party inference hosts including fal.ai, so you can hit it via API without going through MiniMax's own platform.

The audience here is developers and studios building video generation into a product, not casual users looking for a web app to play with — this is a model release first. Open weights are the strategic differentiator: if you can run it yourself, you can fine-tune on your own footage, keep client material off a third-party server, and avoid per-generation API costs at volume. That's a meaningfully different proposition from Veo 3 or Kling, both of which are closed. The tradeoff is that a 2K video model with audio is not something you casually self-host, and MiniMax's tooling and documentation remain thinner than what US labs ship.

Key features

Omni-reference inputs

Accepts text, image, video and audio references in a single prompt and reconciles them, so you can lock a character's appearance from a still while borrowing camera motion from a clip. Most rival models handle one conditioning signal at a time.

Native 2K generation

Outputs at 2K resolution directly rather than generating at 720p or 1080p and upscaling. Fewer artifacts survive into the final frame, and the footage drops into a real edit without an intermediate enhancement step.

Synchronized audio

Generates a matched audio track alongside the video in the same pass instead of leaving sound design as a separate job. Puts H3 in the same category as Veo 3 and ahead of most open video models, which are silent.

Open weights

The model is released with downloadable weights, so teams can self-host, fine-tune on proprietary footage, and keep sensitive material off external APIs — a route that closed models like Veo 3, Kling and Runway simply don't offer.

Pricing

Free tier: Yes — a limited free tier is available for evaluation. Check minimax.io for current generation limits.

Free tier $0

Limited trial generations through MiniMax's platform to evaluate output quality before committing to API spend.

API (token-based) Usage-based

Pay-per-generation pricing through MiniMax's platform and third-party hosts such as fal.ai. Rates vary by resolution and clip length — check the official pricing page for current numbers.

Self-hosted Your own compute

Open weights can be downloaded and run on your own GPUs with no per-generation fee. Hardware requirements for 2K video with audio are substantial.

Pros & cons

Pros

  • Open weights make self-hosting and fine-tuning possible — rare for a video model at this capability level
  • Combines text, image, video and audio references in one prompt instead of forcing a single conditioning mode
  • Native 2K output skips the upscaling step most generators still require
  • Generates synchronized audio in the same pass, not as a separate job
  • Already served by third-party inference hosts like fal.ai, so you can test it without onboarding to MiniMax's platform

Cons

  • ×The roughly 15-second ceiling means anything longer has to be stitched from multiple generations, with the usual continuity problems
  • ×Self-hosting a 2K video model with audio needs serious GPU capacity — the open weights are practically out of reach for individuals
  • ×Token-based API pricing makes per-project cost hard to forecast compared to a flat monthly plan
  • ×Documentation and developer tooling are thinner than Google's or Runway's, and Chinese-lab origin raises compliance questions for some enterprise buyers

How it compares

Compare head-to-head

Comparison explorer Put MiniMax H3 up against any three tools Opens with MiniMax H3 already loaded. Add up to three more from the full index and read pricing, features, pros and cons in one table.

Related reading

Ready to try MiniMax H3?

Head to the official site to start with MiniMax H3 — pricing and plans are listed above.

Visit MiniMax H3
← More Video tools