The Next App Boom Opportunity: When Real-Time Video Generation Arrives
The Next App Explosion: When Real-Time Video Generation Arrives
The Next Scene Is Decided by Your Idea

👦🏻 Author: Aki
🧑🎨 Layout: NCon

For the past two years, AI video competed on generation quality. Starting two weeks ago, it competes on inference speed and real-time performance. Real-time interaction built on post-trained versions of the MiniMax H3 video model, plus generational leaps from GPT-6 and Fable 5.1, all collided in the same week. A new opportunity is opening up for AI applications and models: video is no longer the end product of content — it's becoming the infrastructure of interactive entertainment.
1. New models keep emerging, each ruling for about two weeks
The shrinking-circle war at the model layer is still accelerating. Anthropic and OpenAI successively released Claude Fable 5.1 and GPT-6 Astra, with GPT-6 officially described as a generational leap in computer use and software engineering — RSI and AGI suddenly don't seem so far away. Thanks to these capabilities, people everywhere started experimenting with Blender modeling and 3D game development. AI-generated gameplay is approaching genuinely playable for the first time, and coding and 3D have been folded into new technical pipelines as reference material for video generation.
On the multimodal front, MiniMax open-sourced the H3 weights in early August: a 33B-parameter dense video model with text, image, video, and audio in a single context, generating videos that come with dialogue, sound effects, and music built in. The significance of open-sourcing isn't how strong the model itself is — it's that "post-training" was handed to the entire community. LoRA, distillation, quantization, sparse attention: dozens of teams are simultaneously accelerating and post-training on the same base.
The fastest mover was fal — a leading inference marketplace and infrastructure company. It post-trained H3 first, then moved it onto its own inference engine, producing H3 Max: 5-second 768p clips with audio, generated in 3 seconds, at significantly lower cost.

The community was even faster than fal. Within days, developers built endlessly scrolling short-video feeds, a 24-hour AI news channel fed by RSS and community discussions, real-time Pictionary, dating simulators, and branching text adventures. Most were rough, but they proved one thing: when generation time falls below playback time, "video" is no longer something produced and then watched — it can be decided on the spot.
Tao Zhang, co-founder of Manus, also posted on Jike and WeChat Moments: "The next opportunity for an app explosion: when a Seedance with sub-one-second latency arrives. Originally posted exactly two months ago — emphasizing it again."

That judgment holds. Beyond community developers, some companies have been quietly positioning themselves. AI interactive content company Yoroll just released its own post-trained model, Yoroll H3 Superfast, which infers 10 seconds of video in 4 seconds, and launched YoLive, a real-time narrative product powered by it.
And the company behind it happens to be one of the earliest and most prolific builders in AI interactive games and interactive video.

Before this, the company and the creators on its platform had already racked up a few numbers:
Zombie Scavenger and Republican-Era Mysteries each surpassed 50,000 Steam wishlists; Legend of Huajun: Flip, Your Majesty hit 2 million players within a month of launching on Douyin mini-games; and AI Odyssey, built by a solo creator in two days, once topped the X trending chart in the United States.
2. Faster-than-real-time video exists — but can it be played?
According to the official description, Yoroll H3 Superfast is a post-trained and inference-accelerated version of the open-source MiniMax H3. The official positioning is one sentence: "Built for games, not just for film." The numbers: 10 seconds, 768p, 24fps, with native audio, generated in 4 seconds on 8 B200s — 2.5x real time — running on Yoroll's own inference stack, post-trained on large volumes of game recording data. For comparison, fal's H3 Max produces 5-second clips in under 3 seconds, about 1.7x real time. Superfast is already live on the Yoroll creator platform.
At the same time, Yoroll released a new product built on this model, YoLive (yo.live[1]), with a very direct product form: an interactive narrative game that everyone participates in and that keeps moving forward.

Enter YoLive and viewers see a continuously playing AI interactive narrative channel. Anyone can propose what should happen in the next scene — make the robot dance, stage a rooftop chase, add a twist nobody saw coming — and others vote on the proposals. The winning direction is combined with the established plot, character relationships, current location, and quest state, and continued into new footage, generated on the spot by the H3 Superfast model. The first open channel comes from Zombie Scavenger: the robot cowboy, the model, the ostrich, and that atompunk wasteland form a stable cast and world, within which the audience decides where the story goes.
This gives danmu (bullet comments) a functional role for the first time. People used to chat alongside the same work; now the chat enters the work itself. One person cares about a character's fate, another wants to land a joke, someone else specializes in derailing the story — different intentions collide and become material for the next scene, and a reason to stay in the livestream. For participants, the most compelling moment is "that idea just now was mine."

Real-time generation and livestream gameplay are now live on the Yoroll App and Douyin livestreams
Even more than YoLive, what demonstrates that "real-time generation is a form of gameplay" is the batch of mini-games being incubated by creators on the platform. What they share: gameplay written in Yoroll Code, visuals generated via the H3 Superfast API, and real-time input that directly changes the story. All are still in development and exploration, but their themes are already distinct.

Across these mini-games, "real-time" has already been placed into concrete interaction design. Blind Date Warrior lets players customize a date to their preferences and drive the date through free dialogue and actions: one probing line, one bold move, can push the next scene in a different direction. Streamer Unboxing Simulator turns livestream danmu into gift commands — viewers decide what goes in the box, and the AI generates the unboxing and the streamer's reaction — delight or horror, the players set the question. Infinite Matchstick Adventure lets players freely decide the actions of the matchstick man on paper: switch identities, come up with an unexpected escape, even offer answers outside the preset options, and the AI continues those imaginings into new visuals and story. The most direct change from real-time generation is turning a player's idea into what happens next in the game.

Blind Date Warrior, a UGC real-time game built on the H3 Superfast model
These games are all small in scope, but they're the first batch of works we've seen that treat "the next scene is decided by player input" as the core loop rather than a marketing gimmick.

Yoroll Code's gameplay and level development tools now support GPT-6 Astra and H3 Superfast
3. Why would anyone want to play AI-made content?
Yoroll didn't enter this space because real-time generation arrived. On the contrary, it entered early and built a lot: mini-games validated by player data, large-scale games reaching premium quality, and top creators already onboard. Whether real-time or pre-generated, what users and players care about most is still good stories, fun gameplay, and compelling interaction.
Legend of Huajun: Flip, Your Majesty, made by one of the platform's first creators, is female-oriented interactive content. After launching on Douyin mini-games in August, it reached 2 million players in under a month, making it — as far as we know — one of the most-played AI interactive film-games anywhere.

Legend of Huajun intro and its 2 million players
Overseas, a North American creator spent two days and roughly $250 on Yoroll to make AI Odyssey — a vertical-screen interactive film-game with over 140 video segments and 11 endings, launched on the Yoroll platform and TikTok Minis. Riding the wave of Nolan's The Odyssey film release, this AI Odyssey interactive game once hit #1 on the U.S. X trending chart.

And on large-scale games, Yoroll is taking several flagship AI titles into territory traditionally held by big game studios.

Last week in Cologne, Germany, at Yoroll's booth at gamescom — the world's largest game show — most players who sat down at the computers expected to watch an AI short film. Zombie Scavenger was, after all, originally a widely circulated AI short.
But the demo's structure was far more complex than the Steam page had previously revealed. Beyond the choices and QTEs typical of interactive film-games, there was scene exploration, third-person shooting, and idle tower defense. Players control the robot protagonist clearing zombies on a beach; entering the subway station, the view narrows and the horde swarms in. When the action segment ends, it returns to cinematic narrative, with inventory, weapons, and stats carried into subsequent segments.
In the current 90-minute playable demo, gamescom playtest data showed average player session time exceeding half an hour. The first reaction from overseas players: a game that looks like an AI film can actually let you move through space, explore, and shoot. This shows that AI content can win over hardcore players too.
The gameplay and systems behind these games have all become capabilities of Yoroll's creation platform. As AI video quality keeps rising, gameplay programming keeps improving, and real-time generation keeps accelerating, an explosion of interactive entertainment products is imminent.
4. Open-source models level the starting line
Back to the models, here's an interesting question: a month after H3 went open source, with dozens of teams post-training on the same base, why can companies like fal and Yoroll deliver good results?
Zoom out for a wider view. Today's LLM landscape is "two superpowers, many strong contenders": GPT and Claude lead closed-source, while the strong contenders are almost all open — DeepSeek, Qwen, GLM, Moonshot AI's Kimi K3, and so on. The open-source contenders didn't win the benchmarks, but they won the application layer: lower costs, a more open ecosystem (no account bans or concurrency limits), and the ability to fine-tune a stronger base for your own use cases. Video models are "one superpower, many strong contenders," with most of the contenders being Chinese teams that had almost all stayed closed until now. What happened in the month after H3 released its weights closely resembles what happened in the LLM application layer in the half-year after DeepSeek open-sourced: LoRA, distillation, quantization, sparse attention — multiple versions sprouting from the same base, and the model itself ceasing to be scarce.

The LLM usage leaderboard on OpenRouter over the past month
What open source levels is the starting line at the model layer. What it doesn't level is three things: the inference stack, data, and product.
fal holds the inference stack and the developers. It was an inference infrastructure company to begin with; H3 Max's speed comes partly from post-training and partly from moving the model onto its own inference stack. fal turned "faster than real time" into a public good that anyone can buy by the second.
Yoroll holds data and product. That's also why it needed — and was able — to build its own version.
First, data. From day one of building interactive film-games and 3D games, Yoroll has been accumulating a kind of material nobody else has: in-game footage. First-party long-form content plus works from the creator platform left behind large volumes of footage annotated with branches, states, and gameplay — what kind of shot follows what kind of choice, how a character stays consistent over 90 minutes, how shooting segments and cutscene segments connect. General-purpose video models train on films and short videos; they're good at "shooting," not at "playing." H3 Superfast's post-training uses in-game data to fill that gap: stability of game art styles, character consistency across long sequences, coordination between camera and gameplay pacing, and action control reserved for world models — letting the model continue visuals based on player operations, not just text prompts.
Second, product. Speed only has value when it gets consumed. On the day Superfast launched, it was already being used in three places: generating videos on the creator platform, receiving audience votes in YoLive, and powering real-time interactive mini-games. Every speedup lands directly in players' hands, and player reactions flow back to tell the model team what to fix first — a character failed to execute an action, a prop vanished between shots, the footage doesn't connect. Model-only companies can't get these questions; application-only companies historically couldn't get the model. Open source solved the latter half; the former half can only come from building your own consumer products.
Finally, the inference stack. The bigger challenge in real-time generation is actually inference optimization and acceleration. Anthropic's reported $6 billion bid to acquire Decart (an AI interactive video and real-time generation company) is largely about valuing exactly this capability and accumulation within the team. We understand that Yoroll's parent company LinearGame is in talks to invest in an AI inference chip company — deep collaboration with co-optimized software and hardware, using dedicated "cloud + edge" chips to further realize real-time inference generation and solve the speed and cost problems.

Open-source bases will be replaced generation after generation; the only things that can migrate with them are data, product, and the inference stack. These three are what can genuinely be accumulated in this cycle.
5. The future belongs to "pre-generation + real-time generation"
Pre-generation handles polish; real-time generation handles response. Main storylines, characters, and key scenes are produced in advance; players' unpreset questions and actions are picked up by real-time generation. Rule and state systems remember what happened, keeping story and gameplay coherent.
Cost will determine how the two divide the work: real-time streams watched and shaped by many people together split costs more easily; personalized interaction requires willingness to pay; and freely explorable persistent worlds await further improvements in control capability and inference cost.
Works can also grow while being played: deliver the main storyline first, then polish the characters, spaces, and interactions players care about into follow-up content.
Long-term competitiveness comes from content and IP, interaction data, creation tools, distribution capability, and inference cost control. Yoroll is building in all of these directions. But ultimately, a work still has to answer: Is it fun? Who's willing to spend time on it? Who's willing to pay?
The next scene starts with one of your danmu choices.


References
[1] yo.live: http://yo.live/