← Models

Realtime video generation

MiniMax FastH3

Video on main_video and generated audio downmixed to mono on main_audio, aligned per window. Served on Windflow as a realtime session you can plug into your product and pay for by the second.

Overview

What MiniMax FastH3 does

MiniMax FastH3 is a 35B MiniMax H3 text-to-video-and-audio model, DMD2-distilled to four DiT calls with 90% sparse VSA-H3 attention. On Windflow it runs as a realtime session: connect once, steer it with text prompts while it runs, and receive the live stream back on a named output track.

Every model in the catalog speaks the same typed session protocol, so switching between models changes the controls you send, not how you integrate.

Architecture
35B MiniMax H3 text-to-video-and-audio model, DMD2-distilled to four DiT calls with 90% sparse VSA-H3 attention
Resolution
960 × 544 on the current 4 × H100 realtime profile
Playout
24
Chunk
345 frames per window (14.38 s at 24 fps; published as the 15 s profile)
Control latency
A whole window. The pipeline materialises the clip before delivery, so time-to-first-frame equals one generation.
Inputs
A text prompt with visual, dialogue, soundscape, and music direction
Outputs
Video on main_video and generated audio downmixed to mono on main_audio, aligned per window
Delivery
WebRTC
Session protocol
sdk.realtime.v2
Price
$0.0150 per ready second

$0.0150

Per ready second

Billed only while the session is ready

$54.00

Per ready hour

The same rate, held for an hour

Preview

Runtime status

Capacity is brought up on demand

Private beta

Access

Request access or a live demo

Run MiniMax FastH3 inside your product.