MiniMax-H3

MiniMax-H3

MiniMax’s next-generation native audio-video integrated large model
2026-07-31
Video Generation
Pricing:
$0.0715 /Secondstarting from
Bulk order? Contact your manager for exclusive deals

API Overview

MiniMax-H3 is MiniMax’s next-generation “natively integrated audio-video” video-generation large model, specially designed for commercial-grade film and television production and advertising marketing. Its core selling point lies in breaking the traditional AI video limitations of “visuals without sound” and ultra-short duration. It natively supports 2K high-definition image quality, multi-track stereo sound synthesis, and groundbreakingly extends the continuous video generation length per session to up to 15 seconds.

───────────────────────────────────────────────────────────────────

Core Capabilities

Natively integrated 2K audio-visual experience with ultra-long 15-second continuous generation: This breaks the traditional AI video limitation of 4-5 second clips. The model natively supports generating up to 15 seconds of continuous, long-form video in one go, with image resolution reaching up to 2K level. Built-in native audio synthesis operators enable the model to automatically match and generate synchronized stereo sound effects while producing high-fidelity animated visuals.

Start-and-end frame control and multi-reference image capability: The model supports multi-dimensional reference inputs including images, videos, and audio. Developers only need to upload the start and end frames, and the model will automatically fill in the intermediate frames with logically coherent camera movements and subject transformation processes. When multiple images or character references are provided, the model can precisely maintain high-fidelity consistency in the appearance and visual style of the same person or specific subject.

Precise instruction editing and brand text rendering: This solves the common issue of “garbled text” in traditional image/video generation models. Based on natural language instructions, the model can accurately render clear and readable commercial brand names, packaging subtitles, or UI elements within the generated dynamic visuals. It also supports instruction-level retouching and motion transfer for specific local areas.

Ready-to-use and high-concurrency, low-latency response: Optimized for enterprise-level high-frequency marketing material pipelines through operator acceleration and cloud-based computing power scheduling. The video inference and rendering speed has significantly improved compared to previous generations, greatly reducing the waiting time for high-resolution video rendering.

API Console

Log in to explore more features! Click to Log In

API Analytics

API Reference (2)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
Video(MiniMax-H3)
POST
Stable
View Details
Query(Result)
GET
Stable
View Details

API Pricing

$
ModelDescription302.AI Price

MiniMax-H3

Image / Text-to-Video 768P Pricing rules for requests with reference images: Free for up to 5 images $0.0285 per image for any exceeding 5 images

$0.0715 / Second

MiniMax-H3

Image / Text-to-Video 2K Pricing rules for requests with reference images: Free for up to 5 images $0.0285 per image for any exceeding 5 images

$0.115 / Second

Query

Task Query

Free