Aakib Ansari.
Back to articles
News Brief

MiniMax H3 Generates 2K Video With Native Stereo Audio in One Shot

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
2 min readModel: MiniMax H3
MiniMax H3 Generates 2K Video With Native Stereo Audio in One Shot

MiniMax launched H3 on July 31 — also distributed as Hailuo H3 or Hailuo 3.0 — and the meaningful part isn't resolution or clip length. It's that the model generates synchronised stereo audio alongside the video in a single inference pass, without a separate audio model or a post-processing pipeline.

Previous MiniMax / Hailuo models were capable video generators but specialised tools. H3 is built differently: it takes text, images, video, and audio as simultaneous inputs and generates across those modalities in one context. In practice, you can give it a reference image, a script, and an audio sample and ask for a scene that incorporates all three — rather than generating video first and audio-matching it separately.

The output spec: 2560 × 1440 (2K) at 24fps, clips from 5 to 15 seconds. MiniMax reports 2K generation pricing significantly lower than "mainstream industry alternatives" — specific numbers haven't been published in the launch materials, but the API pricing is live and listed on the MiniMax Open Platform.

Omni-referencing is the other notable feature: the model can take multiple reference images, videos, and audio tracks simultaneously and use them to anchor style, character, or environment across a generated clip. Combined with instruction-based editing — swapping a character or background via natural language without re-generating the whole scene — it positions H3 toward commercial creative workflows: advertising, e-commerce product visualisation, and film pre-vis.

Access is available now via the Hailuo AI app and the MiniMax Open Platform API. Open weights are committed for release "in the coming days" following the July 31 launch, though no exact date has been confirmed. MiniMax has not disclosed the model's parameter count or architecture in the launch materials.

The model should not be confused with Kling O3, a separate video generation model from Kuaishou released in the same period.

Related Articles

Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model
News Brief2 min read
Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model

Meta Superintelligence Labs has released Muse Glimmer, a 30-billion parameter open-weight model distilled from its proprietary Muse Spark flagship. Published under an Apache 2.0 license, Glimmer is purpose-built for offline, on-device agentic workloads like coding, debugging, and file management on consumer hardware.

Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License
News Brief2 min read
Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License

Ant Group's inclusionAI team has released Ling 3.0 Flash FP8, a highly efficient 124-billion parameter Mixture-of-Experts (MoE) model. Featuring an MIT license and a custom hybrid attention architecture, the model reduces active parameters to 5.1 billion per token, matching the performance of much larger models while dramatically lowering operational costs.

Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks
News Brief2 min read
Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks

Liquid AI released LFM2.5-2.6B on August 4, a 2.69B-parameter on-device model purpose-built for agentic workloads. Using a hybrid architecture of short convolution blocks and grouped query attention, it runs under 2.5 GB of memory and reaches approximately 220 tokens/s on Apple M5 Max — while matching or exceeding Qwen3.5-9B on tool use and instruction-following benchmarks.