World models: the kids' table for those who couldn't make it to the real game
Sitting at the kids' table
"Sitting at the Kids' Table"
It's September 2026, and we can basically declare that "world models" is a garbage concept.
World models originated from a bunch of old guys who couldn't get a seat at the LLM table, trying to force their way in, failing to compete in large language models—which had already become heavy industry, cutthroat manufacturing—and so inventing an entirely new "world model" track.
The figureheads of world models, Fei-Fei Li and Yann LeCun, are working on completely different things. Grouping fundamentally different things under one concept is already pretty funny.
Even funnier: a swarm of domestic prodigies immediately popped up to claim they were building world models too.
Of course, these prodigies all have excellent communication skills. Bullshitting is their core moat. Ask them what a world model is, and they'll inevitably ramble from their family of origin through technical architecture to their convoluted academic lineage with various big names.
Bottom line: "Prodigies, world models, send money."
The reality, however: Post-training open-source video models. No compute, no data, no users. Revenue entirely from VC arbitrage.
And these prodigies all follow the exact same playbook. After raising money, they immediately rush to AgiBot, Synced, and QbitAI for publicity. Their only innovation is adjectival innovation—stacking modifiers before "world model."
Let me briefly list some:
Causal world model, social world model, societal world model, cognitive world model, neurocognitive world model, memory-augmented world model, brain-inspired world action model, psychological world model, mental world model, subjective world model, two-layer world model, dual-agent world model, human social world model, physical-social unified world model, continuously self-evolving enterprise world model; plus Jiemian world model, AI4AI world model, world synesthesia model, world unified model, recurrent world model, enlightenment world model, molecular world model, micro-molecular world model, microscopic world model, hour-level world model, Flash World Model...
And then there's this XGEN that just dropped yesterday. On the surface: "Fuck your world model." In reality: "We're building a subjective experience model."
Holy shit, you just changed the name, and swapped in an even more pretentious term to annoy us.

From my observation, there are four main types of people pushing world model narratives:
Type one: Application companies that couldn't hack it, pivoting on the spot to Neo labs, trying to squeeze more money out of model stories. Typical example: Loopit. AI Douyin pivoting to world models—and people actually sat through that pitch.
Type two: Video model companies that couldn't hack it, post-training a real-time generation video model and calling it a world model. Typical examples: Shengshu Technology, AISphere. Though to be fair, if these real-time generation models focused on specific use cases, they might actually be useful.
Type three is the best: AI 3D model companies pivoting to world models. Tripo, Yingmo, Meshy—all raised serious money this way. This is actually good.
Type four is the worst: PhD students dropping out of school, quitting jobs, egged on by investors desperate to get in on the action. Whatever's hot, they do it. Cranking out digital humans and janky GTA-style short videos, then daring to call it world models.
Setting aside the potentially valuable AI 3D models—those guys are just trying to tell stories and raise money, which is understandable—the rest are all post-training video models.
The only progress these past few months: when they were fundraising, everyone could only post-train on Wan and LTX. Both models' newer versions have been open-sourced for nearly a year, and the results are pretty rough.
Fortunately, Junjie Yan—our sharp-browed, starry-eyed savior from the heavens—threw these prodigies a lifeline. MiniMax H3 is the only open-source video model that can go toe-to-toe with Seedance 2.0.
So you can see these world model companies starting to密集宣发 products lately, purely because post-training on H3 yields better results, which means another fundraising cycle.
This creates a pretty hilarious dynamic.
MiniMax is a large model company whose foundation model capabilities have fallen behind. Though I firmly believe it's fundamentally no different from Zhipu AI or Moonshot AI—just lagging by one release, temporarily behind. But its valuation reflects market consensus: market cap at only 1/3 to 1/4 of Zhipu AI's.
The MiniMax that large model companies mostly look down on is the undisputed big daddy of world models.
I sincerely suggest that these world model companies, after scamming their funding, donate half their raised capital to Junjie Yan, with a note saying "voluntary gift." It's all fundamentally post-training on H3. As long as MiniMax stays alive, you can keep pumping the hype indefinitely.
Junjie Yan is the one and only sun of world models. Whatever trickles from between his generous fingers is enough for this swarm of bullshitting prodigies to squeeze another two rounds of investor cash.
What does this tell us?
The most badass world model is still under the table. Second-tier large models are still above the table. The boundary is crystal clear—only those who can't get a seat go build world models.
The logic is simple. Large models are manufacturing with converged technical approaches. The biggest difference between teams is compute and data.
A large model company CEO has two critical jobs: securing compute, and securing data. And I firmly believe MiniMax hasn't been knocked off the table because its compute and data aren't significantly behind—so there's no reason it should stay down long-term.
Of course, this assumption excludes Zcode secretly uploading user code data. I'm absolutely certain Zhipu AI would never use that user data to train models!
Back to world models.
These prodigies all claim they're training models too, raising money at model company valuations. But the biggest problem: most have no compute, no data, can only struggle to do some post-training, and their products have zero differentiation:
The entire market is real-time generated digital human livestreams and WASD game footage, all roughly the same quality, and all these people claim their costs are an order of magnitude lower than competitors.
Despair, friends. Pure平地干拔. These things are first of all useless, and second of all undifferentiated. I feel extreme scarcity, deeply nostalgic for the Macaron and Leon Ming who could actually tell innovative stories.
Another thought: there's just too much hot money, investors are truly terrible, and everyone is just desperate for outsized returns.
The previous narrative cycle: investors egged on product managers to start companies, resulting in a pile of useless products nobody used.
This narrative cycle: investors egged on researchers to start companies, resulting in a pile of useless post-training video models.
Besides these investors getting promoted off their deals, and the occasional founder cashing out secondaries, everyone knows we're all half-hearted, yet we all pretend to seriously discuss the technical questions of this赛道.
I laugh whenever I hear people discussing how many technical approaches world models have. This isn't a technical question. It's purely a question of motives.
Investors' motive for backing world models: AI application stories have temporarily cooled, large models are so hot, but DeepSeek and Moonshot AI valuations are too high to get in. World models are a large model concept stock—invest in them, easy to justify, and supposedly building models means you can't be falsified in the short term.
The prodigies' motive for building world models is even simpler. "Sheng Yang was nothing special back then, and Junjie Yan started with metaverse digital humans. If they can build models, so can I."
Timing, timing! It's all about timing.
Large models are pure manufacturing—labor-intensive heavy industry. Every model shop brags about their tens of thousands of self-built GPU clusters, their unique data pipelines.
This industry is so heavy that Zhipu AI has raised nearly $10 billion in the secondary market just this half-year. You prodigies plan to compete on compute and data with what exactly?
Of course, if a prodigy spent their entire funding on buying B300s, actually acquired them, and considering the terrifying growth rate of compute demand—investing in these world models might actually make money.
So yeah, the most important thing in building world models is don't innovate. Spend all the money on compute. Guaranteed returns, bro.
At this moment, I can't help but miss the conceptually most advanced startup in real-time generation: our dear Liu Yu's Vivix.
Vivix has the most legitimate claim to the world model story. They were ahead of the curve by over a year, already pushing real-time video generation narratives in early 2025. Cash on hand higher than many world model teams' valuations.
Yet Vivix doesn't call what they do world models. Setting aside actual results, they're honest and grounded, objectively describing their work as streaming video models. Once when Zang AI accidentally described Vivix as a world model, they specifically asked us to correct it.
All I can say is, amid this insane world model frenzy, Vivix and Brother Yu are in any case hardworking, grounded, and don't bullshit arbitrarily. My heart is full of gratitude ❤️
Also, a heads-up: we've prepared a livestream Bench, head-to-head testing these real-time video generation models' capabilities, to see which H3 post-training version is strongest, or if any hero has surpassed H3. Stay tuned!
(This article's cover image was generated by ChatGPT; purely human-written.)
Subscribe to our Substack at funeralai.substack.com