Realtime video generation
MiniMax FastH3
Video on main_video and generated audio downmixed to mono on main_audio, aligned per window. Served on Windflow as a realtime session you can plug into your product and pay for by the second.
Overview
What MiniMax FastH3 does
MiniMax FastH3 is a 35B MiniMax H3 text-to-video-and-audio model, DMD2-distilled to four DiT calls with 90% sparse VSA-H3 attention. On Windflow it runs as a realtime session: connect once, steer it with text prompts while it runs, and receive the live stream back on a named output track.
Every model in the catalog speaks the same typed session protocol, so switching between models changes the controls you send, not how you integrate.
- Architecture
- 35B MiniMax H3 text-to-video-and-audio model, DMD2-distilled to four DiT calls with 90% sparse VSA-H3 attention
- Resolution
- 960 × 544 on the current 4 × H100 realtime profile
- Playout
- 24
- Chunk
- 345 frames per window (14.38 s at 24 fps; published as the 15 s profile)
- Control latency
- A whole window. The pipeline materialises the clip before delivery, so time-to-first-frame equals one generation.
- Inputs
- A text prompt with visual, dialogue, soundscape, and music direction
- Outputs
- Video on main_video and generated audio downmixed to mono on main_audio, aligned per window
- Delivery
- WebRTC
- Session protocol
- sdk.realtime.v2
- Price
- $0.0150 per ready second
$0.0150
Per ready second
Billed only while the session is ready
$54.00
Per ready hour
The same rate, held for an hour
Preview
Runtime status
Capacity is brought up on demand
Private beta
Access
Request access or a live demo
Other models
