MiniMax H3 Max: Faster Than Playback, AI Video Enters the Era of Real-Time Interaction | Luzhou Shengmingli
24-Hour AI Livestream Launches

Oasis Capital portfolio company MiniMax recently released its H3 Max model: generating a full 5-second, 768p audio-video clip takes under 3 seconds, enabling video generation to "outpace playback" for the first time and cross the critical speed threshold for real-time interaction.**
Since H3's open-source release over three weeks ago, it has been downloaded more than 24 million times and spawned over 300 publicly available derivative models. Developers have already built real-time livestreams and 24-hour AI TV stations on top of H3 Max, as video models begin evolving from content generation tools into continuously running content systems.
With advances in complex instruction understanding, real-time interaction, and multimodal content generation, H3 Max provides a more stable and efficient technical foundation for applications like livestreaming, gaming, and interactive storytelling. It also signals that large models are accelerating their integration into real business operations and everyday experiences, unlocking broader application value.
Below is MiniMax's full share. Enjoy.
Fal engineer Alex Koumpas turned his Twitch channel into a writers' room.
Someone typed a prompt asking for a street scene cut; the next viewer wanted a 1960s-style puppet commercial; another demanded the frame suddenly shift to a robot face tangled in countless wires. The chat flooded with ideas in English, Spanish, Japanese, Russian, Chinese — suggestions from users around the world.
Koumpas gave the stream a straightforward title: "Chat Directs the Show with MiniMax H3 Max."
Twitch — koumpas channel (video sped up) This stream is a microcosm of how quickly the developer ecosystem has grown since H3's open-source release.
After open-sourcing H3, we continued training and optimizing the model with ecosystem partners around different use cases, making these variants available to the community:
- Partnered with Running Hub to create MiniMax H3 480P for multi-scenario needs
- Worked with fal on post-training and inference optimization for real-time generation, launching H3 Max 768P and H3 Max 480P
- And collaborated with FastVideo on FastH3, using sparsification and hardware adaptation
Today, we're bringing H3 Max 768P and H3 Max 480P to the Open Platform and MiniMax Design.
H3 Max generates a 5-second, 768p audio-video clip in under 3 seconds, crossing the speed threshold required for real-time livestreaming. While one clip is still playing, the system already has time to prepare the next.
Overseas developers have already built Twitch livestreams, 24-hour "AI TV stations," and creative experiments inviting viewers to rewrite plots in real time on top of H3 Max. Video models are being plugged into continuously running content streams, beginning to participate in real-time interaction where every viewer command can change what appears next.

3-Second Generation, Real-Time Co-Creation with Viewers
In this wave of AI livestreaming, fal engineer Rehan Sheikh's Twitch stream was the first to go viral. Rehan connected H3 Max directly into a livestream feed, creating a "trans-dimensional TV" with no fixed programming schedule. The model generated new visuals and audio nonstop within seconds, fed directly into the live broadcast. The concept has racked up nearly 3.5 million views on X.

Soon after, prominent developer Pieter Levels launched a 24-hour AI livestream site. The homepage had just one simple rule: "The chat decides what airs next."
Viewers typed into the chat, and the AI attempted to connect new requests with the previous scene before generating the next clip. In this micro TV station with only seconds of inventory, the broadcast might suddenly cut from animation to a live interview, then jump into some absurdist commercial. The audience's next message would send the story in another direction entirely.

Unlike pre-recording a demo video for broadcast, real-time livestreaming requires simultaneously handling unfamiliar prompts, video generation, queue sequencing, and continuous streaming. The model's speed, instruction comprehension, and audio-video capabilities are all tested in the same public pipeline, with every command directly observable: viewer issues command, model generates clip, clip enters livestream, new command follows.
This is precisely why generation speed determines whether the livestream can sustain itself.

Generation Faster Than Playback, Unlocking Instant Livestreaming Scenarios
H3 Max can support this kind of livestreaming because it's fast enough.
H3 Max supports text-to-video and image-to-video, generating complete audio-video at 480p or 768p, 5 to 15 seconds, at 24fps. A 5-second, 768p video generates in under 3 seconds; a 15-second video completes in roughly 15 seconds. Through coordinated optimization of post-training and inference engines, H3 Max achieves approximately 35x the throughput of the base MiniMax H3, ranking first on Artificial Analysis and Design Arena's image-to-video leaderboards.

Building on MiniMax H3's open-source foundation, fal added new data and verifiable reinforcement learning to continuously optimize prompt adherence and visual quality, while adapting to its own inference infrastructure for targeted real-time generation optimization. A 5-second video generates in under 3 seconds, theoretically leaving headroom for queuing, content moderation, encoding, and streaming. While real-time livestreaming remains subject to concurrency and network conditions, with occasional stuttering, H3 Max already meets the baseline speed requirements for this category of real-time experiments.
Ancient Greek vase painting style consistency demo — fal official website
The FastVideo team took a different route in adapting H3. FastVideo, Nuva Lab, and NVIDIA FastGen jointly released FastH3, compressing the base H3's 49 Transformer forward passes down to 4, combined with 90% sparsity VSA acceleration.
FastH3 uses 8 B200s to generate a 15-second, 768p video in 13 seconds; on a single Blackwell GPU, it achieves up to 14x speedup over the base H3 running on equivalent hardware. FastH3 has publicly released model weights, LoRA, and inference solutions, with the current version supporting text-to-audio-video generation.
FastH3 — Hao AI Lab official demo

Open Source and Openness, Letting Creativity Emerge
We're excited to see new forms of video generation and interaction continually emerging in the community, built on the open-source H3 model. For MiniMax, the open-source community isn't just users of H3 — they're also critical participants in the model's continued evolution and the ongoing renewal of real-world application scenarios.
Today's 24-hour AI livestreams remain in early stages; concurrency and long-duration operation still need further validation. Voice stability for characters, whether visuals can remain consistent over extended periods, whether inference interfaces suit sustained calling — this real-world feedback from livestreaming scenarios will feed back into the next round of model and system optimization.
What matters is that in this wave of AI livestreaming, video generation is becoming more than a one-off content tool. It's beginning to function as a content system that responds in real time and grows continuously. MiniMax and the open-source community are pushing together on the speed, interactivity, and application boundaries of video models.
As models continue iterating, we believe new interactive content, new creator tools, and products people haven't yet conceived or defined will emerge from the community. We look forward to developers and creators bringing H3 into more concrete scenarios, exploring possibilities that were previously impractical or even unimaginable.
Since H3's open-source release over three weeks ago, downloads have exceeded 24 million, with over 300 publicly available derivative models, making it the most-downloaded model globally in 2026. We will continue exploring alongside the open-source community and ecosystem partners, lowering barriers for developers and enterprises, bringing H3 into more real commercial applications, and solving problems that were previously intractable.
API access for H3 Max: platform.minimaxi.com/docs/api-reference/video-generation-v2-create
MiniMax Design download: design.minimaxi.com





