Ant Group Treats World Models Like a Game
A Flash of Insight on a Hot Stove
"A Flash of Inspiration in the Hot Seat"
One of the biggest headaches in the world-model space is how loosely defined the term is. There are at least three or four factions, each doing their own thing, shipping a product, and declaring theirs the real world model. No one can really prove anyone else wrong.
Lingbo Tech, an Ant Group subsidiary, seems to have solved this problem by simply building in almost every existing direction.
I clicked through each one: VLA 2.0, Vision 1.0, and VA 2.0 are physics-AI systems that help robots understand the world, similar to what NVIDIA's Cosmos is doing; Video 1.0 is a video generation model, but purpose-built for embodied intelligence with an emphasis on physical interaction — feels like a competitor to Sand.ai; World 2.0 is the straightforward one, open-world generation, in the same ballpark as Genie 3, Happy Oyster, and PixVerse R1 that I've written about before.
What's the point of so many models? Is this the seven swords combining to form a Gundam, or just not putting all your eggs in one basket? Hard to say for now.
Some of you must be wondering: how do I even know Ant dropped this many updates? Seems pretty niche, right?
The answer is I've been a loyal user of Lingguang, and Ant stuffed their world model right into it.
That's right, you read that correctly. Not in a car, not in a robot brain — a world model, in a mobile app.
So Lingguang users are all replicants now? They're not even pretending anymore.
I thought Lingguang was done for, but somehow they pulled this Hail Mary of integrating a world model in-place to stay alive. AI-generated content becomes real-time interaction — they're basically telling Vivix's story now.
Product team: "We eating again, colleagues!"
Curious who this "Teacher Sun" is in Lingguang's world model announcement post, though.
Despite the name, this "world model" feature in Lingguang is essentially one-sentence generation of 4399-style minigames — a 3D version of Macaron.
I'm not even hating for no reason. I clicked into a few of Lingguang's preset open worlds and found two Honor of Kings-style mobile control pads added to the bottom left and right corners. Left side moves the character, right side casts skills.
The skill names are peak anime cringe. For example, in a world called "Summer Bliss at the Seaside Convenience Store," the protagonist can unleash "Spacetime · Fission," "Blooming Storm Descends," and "Crystal Realm · Temporal Rain" — which, when you tap the buttons, produce extremely flashy and colorful visual effects directly in front of the character.
But here's the thing: these are purely visual effects. They have no lasting impact on the environment, no effect on how the world develops. Just pure VFX that flash-blind you and vanish, leaving not a cloud behind.
Honestly, this feature is worse than Kingsoft PaintKing I played in the computer lab as a kid. PaintKing had a "Fairy Bag" feature where one mouse click would make flowers, butterflies, and little stars fly across the screen. Also very flashy ↓
I propose Kingsoft PaintKing directly rebrand itself as the mother of all world models.
Back to Lingguang: beyond entering their preset worlds, you can also upload an image to create your own world.
For instance, I sent a Backrooms image, and then ↓
First off, I have to say: those hype-man bloggers who immediately test world models with Backrooms, dreamscapes, and magic worlds whenever they get the chance? Smart move.
Because these places have no logic and no physics rules — they share a certain kinship with this world model's weaknesses.
When clipping through walls, clipping through models, misalignment, and instability inevitably happen, these guys can still thumbs-up and praise: "So creative! That's exactly the Backrooms vibe!"
Second, in Lingguang's preset worlds, character skills get personalized to match each world.
But if users generate their own world, no matter what the scene, you only get fire and lightning.
Commit to the bit, would you? Is adding an "auto-generate matching skills and icons" workflow really that hard?
Or does Lingguang not believe users will stick around long enough to reach the upload-image-and-generate-world stage 😭
Of course, this is a mobile world model embedded inside Lingguang, so there may be performance compression. So I also tried LingBot-World 2.0 on PC.
They put it up on Reactor
And found that the PC version of LingBot-World 2.0 does indeed personalize character actions based on different worlds.
For example, when I generated a "Shenzhen Electronics Factory Worker World," the system automatically assigned this worker five actions: "Welding Task," "Component Sorting," "Machine Startup," "Quality Inspection," and "Tool Reset."
Unfortunately you can't input custom commands, otherwise I really wanted to try making Grandpa Worker sit down and rest.
Anyway, I channeled my inner capitalist, pressed a few buttons while the world model was running, and the result is below ↓
Actually looks pretty decent.
Never mind whether the cyber-worker is actually interacting with a virtual production line — share this video around, and regular folks will absolutely believe Ant is about to stomp Unitree and punch out Boston Dynamics.
But thinking more carefully, I suspect this is just stitched together.
Those previous world models focused on real-time generated video, like Happy Oyster and PixVerse R1, generally had two modes: exploration mode, which generates a 3D space where you move a character around with WASD, and director mode, which generates a scene where you change character states by inputting prompts.
Isn't LingBot-World 2.0 just the "input prompt and hit enter" step macro'd into a hotkey? Doesn't that just look like the character is interacting?
Gotta admit, pretty thoughtful move — directly raising the hype threshold for world models to a whole new level.
Later I did some research and found that adding buttons to world models isn't exactly Ant's original stroke of genius either.
Happy Oyster officially launched last month, and the biggest change from testing was also adding a bunch of buttons.
But no matter the scene or character, Happy Oyster's default actions are always crouch, jump, sprint, and attack.
In the factory not working but hitting coworkers, flying through dreams and still punching the air.
Extremely aggressive. Recommend Freudian therapy.
Beyond that, LingBot-World 2.0 has another serious flaw: avoidant personality disorder.
Previous world models, like Genie 3, when facing walls, trees, and other solid objects, would occasionally clip through them in physics-defying ways, like below ↓
LingBot-World 2.0 doesn't have this problem.
Because when the character walks through Ant's world model, it directly channels Moses: the landscape ahead parts like the Red Sea, and pedestrians and vehicles all actively flee the scene.
Have to admire the LingBot-World 2.0 researchers — sometimes the clipping is so smooth, I can't help but wonder if it was designed this way?
This distortion of human cognition is its own kind of miracle.
More unexpectedly, the best performer on clipping was actually Happy Oyster: not only refusing to clip through walls, but in driving scenes, you can get out of the car and enter buildings ↓


Somewhat ahead of the pack. Feels like if Alibaba doesn't get brainwashed by this market wave of "world model = game demo" heresy, they might actually make something real.
But that attack button at the bottom of the screen and the quest log in the top left are giving me the chills 😰

Actually I get why they're doing this:
Vivix, Vidu, and those Liu-style video generation shops have already told their story well in the digital human livestream track;
We open-world folks also need to hitch ourselves to an existing track to seem useful, and after some thought, AI gaming it is.
Including that recently hyped "three girls hand-crafting an otome game PV" — wasn't that just PixVerse adding dialogue boxes and options to their world model and rebranding it PixVerse·Game.

Haven't gotten an invite code for this one yet, will check it out in a couple days
As you can see, adding buttons to world models, making little people do magic in them, pretending your open world is an open-world game — this has quietly become a new self-indulgent direction for world model companies.
Why self-indulgent? Because from a player's perspective, even if these world models became fully realized, I'd have zero desire to use them.
The reason is simple: as games, they're boring as hell.
Really, even if these world models lasted a full hour, with every asset photorealistic. But with no levels or objectives, no story or character relationships — this isn't a game, it's a walking simulator.
Who plays games just to walk around an open world? If that's your hobby, just go downstairs.
Right? Even widely recognized open-world games that let you "do whatever you want," like GTA5 and The Legend of Zelda: Breath of the Wild, have clear mechanics and storylines. They don't just drop you in Los Santos and say explore wherever you want. That would be as boring as real life.
You play games for either the mechanics or the story. Without either, you're not a game. Fable 5's hand-crafted Tetris is more playable than you.
And this so-called "speak and it happens," "generate whatever you want" is meaningless to players.
Players are people who play games, not people who make games. Players want to directly consume fun served to them on a silver platter, not become fun manufacturers themselves.
At the end of the day, most ordinary users simply don't have content creation talent. You give them unlimited power to generate whatever they want, and they can't even articulate what they want. No wonder Rousseau said man is born free yet everywhere he is in chains 😭
A product without user mindshare is just mud, no matter how strong its capabilities. Fukuyama once made a state capacity vs. accountability two-by-two: the gist being some countries have sound institutions and strong execution — all good, like the Nordics; while others have complete institutions but terrible execution, ending up a mess, like Nigeria.
We really need a coordinate system tailored for AI products: x-axis is product capability, y-axis is user mindshare. Most world model products today (specifically those hard-hitching to gaming) have terrible product capability and zero user mindshare. No one uses them except content creators. They're not even Nigeria — pure Somalias of the AI product world.

Terrifying. All I can say is, whoever wants on this pirate ship, be my guest.
Anyway, though I don't know what these world models are actually good for, I still look forward to more such products emerging and telling increasingly wild stories.
After all, in the dreary AI product scene where everyone's studying how to boost productivity, anyone willing to pour energy and attention into something this uncertain-futured is practically a philanthropist, no?
(Cover image generated by ChatGPT, text purely human-written)
⬇️
Subscribe to our Substack at funeralai.substack.com