To Hell With 'World Models'

Generate unknown.

@Zhiyan Chen

The story begins a very, very long time ago.

Twenty years ago, Che Haoxuan was in third grade, and there was one kid in his class who left a deep impression. The kid was a top student, and every day after school, a family member would pull up in a car to take him home.

Che thought this kid must be happy.

But one day, he started wondering: If I were him, what would the world look like? He grew curious whether the same world might appear completely different through someone else's eyes.

I know what you're thinking: so what?

This is a story Che told me himself. Twenty years later, at age 29, he founded XGEN, hoping to build a "Door of Anywhere" — a portal that transforms worlds once imaginable only in the mind into places you can actually enter.

The concept is a bit convoluted.

"So, is this a world model?" I asked.

"Forget it!" Che replied almost instantly, then threw out something even harder to grasp: interactive experience model, a subjective experience model. He calls what XGEN is building "generative world simulation."

What follows is my attempt to tell you Che Haoxuan's story.

The Guomao Monster, the Police, and the Second World

"What if a monster suddenly appeared in Beijing's Guomao district? What would happen to the world?"

To explain what this Door of Anywhere actually is, Che proposed an outlandish scenario. He said if you asked today's AI to generate a video of this, you'd get a monster climbing buildings, biting people, smashing towers, crowds scattering in panic.

But here's the problem: "The police wouldn't come."

No police, because AI generation doesn't account for the latent rules of this world: someone would pull out a phone — there's a profession called police — police come when called — they show up where people can see them.

So Che wants to generate two worlds.

The first is what the human eye sees. "You wander through a world, and it changes according to your subjective perspective."

The second is what the human eye cannot see. "For example, what's the U.S. President doing right now? Sleeping, or drinking, or something else. I don't know, but he does exist in some state at this moment."

Past visual generation models dealt almost exclusively with the first world. "If you can't see it, it doesn't exist; if it doesn't exist, AI can't model it." By Che's theory, without the second world, even if the monster smashed Guomao for an hour, the police would never appear.

Che wants the second world to exist in another form: states. The unseen world doesn't need every frame drawn out — it only needs to be remembered, to be reasoned through — who is where, what just happened, what memories and goals each entity holds, what changes one action triggers. Only when it enters the field of view does it get rendered into pixels.

Architecturally, the World State layer handles world dynamics and state transitions, while the Render layer converts what the subject should see at this moment into pixels. The two are trained separately but collaborate within the same hierarchical architecture.

There's an industry reference point for this approach: Genie.

Genie is a model family from Google DeepMind, and the most representative "world model" route today: it doesn't generate video for passive viewing, but environments a person can enter, operate within, and that respond to those operations. Genie 2, released in December 2024, can generate interactive environments explorable with keyboard and mouse from a single image. DeepMind claims these generated environments can last up to one minute, though most demos run 10 to 20 seconds.

By Che's account, XGEN represents a modification of the Genie approach. Their current single-point focus is "long-horizon consistency" — not making a few seconds of video look coherent, but attempting to keep characters, states, and causality self-consistent after a world has been running for dozens of minutes.

When I asked how big this could ultimately become, Che said: "People could step into movies, anime, novels, and live alongside characters; robots could learn to do things in virtual worlds and bring those skills back to reality. The internet connected one world; we want everyone to be able to create their own world."

Before founding XGEN, Che had already made one attempt down this path.

From May to September 2024, he interned at Tencent IEG, where as first author he completed GameGen-X. It can generate open-world game video from text, accepting structured text commands and keyboard controls. When players move characters or take actions, the model continues generating environmental and character changes based on the current video segment. Unlike traditional games with pre-made maps and assets, both environment and characters are generated by the model.

The related paper was later accepted by top machine learning conference ICLR 2025. The paper defines GameGen-X as the first diffusion Transformer model to simultaneously support open-world game video generation and interactive control.

GameGen-X was released before Genie 2. When Genie 2 launched, Che wrote on his social feed that Genie 2 and GameGen-X were then the only two works not dependent on existing game worlds, capable of generating new environments. In the same post, he acknowledged that the other team had gone further in control diversity, short-term memory, and behavioral emergence.

I asked him: if the goal is to build unknown worlds, why must there still be a "subjective perspective"?

"To model a fully objective world, you'd need to exhaust all data in the world. I think that's impossible." He drew an analogy to AlphaGo: it didn't exhaust all possible game solutions, but reframed the problem from "find the optimal solution" to "be stronger than the strongest human."

XGEN's emphasis on "subjective perspective" follows the same philosophy.

"If a person is trapped in a room, the signals they can send, the changes they can trigger — those can be covered." Che believes that from a subjective perspective, there are only three possible interactions with the world: manipulation, navigation, and communication.

Thus, subjective perspective draws a boundary around the boundless world: the model doesn't need to simultaneously calculate every person's movements in the city, every room's lighting, every object's pixel changes — it only needs to maintain states relevant to the current subject and their possible actions. When the subject moves, operates, or communicates, relevant portions are called up and unfolded.

Put more simply, XGEN's "subjective experience model" doesn't pursue exhaustive coverage, only plausibility — unseen portions need to be modeled cheaply enough that when they come into view, they don't break the illusion.

Liu Bang, the Stubborn Mule, and Another Possibility

We were sitting in a Western restaurant in Beijing as he said all this.

For the first two hours, Che did almost all the talking. Technology, product, fundraising, organization — the topics kept shifting, and he rarely paused.

Che's eyes are bright, with sharp, upturned outer corners. Because of his delicate features, he was often mistaken for a girl as a child, and even recently someone described him as "a K-pop boy band member dropped into world models."

It's hard to reconcile this person who has answers for everything with the Shaanxi freshman who entered Northwestern Polytechnical University in 2015.

Che initially studied aircraft manufacturing, the university's flagship program. In an atmosphere of "aviation for national service," building airplanes was practically the most prestigious path. The trajectory ahead was clear: Shenyang Aircraft, Chengdu Aircraft, and the major aviation research institutes.

But once in university, he quickly discovered his true fascination lay elsewhere.

"Computers, coding, AI..." He hadn't done any AI projects at that point, couldn't claim systematic study — he just loved it. He described it as a "fateful interest."

He began to sense that while building planes, tanks, and artillery certainly mattered, they felt distant from the problems he actually cared about. He saw poverty, hunger, educational inequality, and disease-induced suffering in the world. Weapons couldn't solve these, "but AI, maybe could."

He recalls his judgment then was still hazy. He couldn't even articulate how AI would address these issues. Just an instinct that AI would be a technology of greater scale — not some particular machine, but an elevation of human knowledge and capability, an iteration in how we process problems.

Also in his first year, he met a classmate from Jiangsu. This student had his own approach to problem-solving: for the final, most difficult math problem on the gaokao, he used university-level methods; when drafting work, he'd divide a sheet into different zones, using it with extreme density.

To Che, these were things he'd never seen before.

"Stunned with admiration."

In retrospect, it was just a method for solving problems and organizing drafts. But it made Che realize there were other methods, other possibilities beyond his immediate environment.

In 2016, AlphaGo defeated Lee Sedol. Faced with nearly inexhaustible variation, the machine could still defeat the strongest human player through learning and search. With the AI seed already planted, he decided to switch majors to software engineering.

Software engineering wasn't particularly popular at NPU at the time, and the tuition was higher. Transferring involved quotas and student resources between colleges — his original college was reluctant to let him go, and the software school rarely accepted transfers. AI was far from being an obvious career path then, and no one could guarantee this road led somewhere better.

After several applications went unanswered, Che decided to write directly to university leadership. He'd read countless strangers' stories on Zhihu, learning one朴素 lesson: for things that truly matter, go directly to the person who can make the decision. His application was eventually approved.

"I have very thin skin." He said he wasn't unafraid of rejection, "but if doing this thing is right, you have to do it."

Similar things kept happening —

After undergrad, Che gave up his guaranteed graduate school spot to pursue a PhD at The Chinese University of Hong Kong. Due to research philosophy conflicts with his advisor, he withdrew and applied to U.S. schools. To go abroad, he raised his TOEFL score from 32 to 102 in six months.

In May 2020, the United States issued Presidential Proclamation 10043. Visa and entry restrictions, compounded by the pandemic, interrupted his original study path, and his newly admitted school became unreachable. During that period, he'd sleep around 4 or 5 a.m. and wake around 4 or 5 p.m. It was then that he posted the first related thread online, organized the first joint letter, and contacted other affected students to appeal to their respective universities for help. Later, they tried petitions and writing to U.S. legislators, calling themselves "little ants."

Blocked from U.S. study, he redirected his applications to The Hong Kong University of Science and Technology. Materials submitted in May 2021, offer received in June, enrolled in July — entering a direction with almost no prior accumulation: model generalization and medical imaging.

He was to answer whether things that ran well in the lab would still work in the real world. Two years later, he reached a discouraging but unsurprising conclusion: "You generalize as far as your data goes."

In February 2024, Sora launched. OpenAI described video generation as a path toward "world simulator." Che saw something bigger again: video might not just be for generating content, but could become an environment where AI gains experience.

So he decided to shift from understanding models to simulating worlds. Three months later, Che interned at Tencent, beginning GameGen.

Che describes himself with three words: gambler, stubborn mule, naive.

"Once you see something bigger, something better, it's actually hard to go back to the slightly lesser thing. You see a better life, better people, better opportunities — don't you want to reach for them?"

At this, Che raised his eyebrows, as if suddenly remembering something, then stitched together the words of Liu Bang and Xiang Yu when they saw Qin Shi Huang's imperial procession: "A great man should be like this" — "He can be replaced."

Childhood, Diablo, and the Door of Anywhere

GameGen hit nearly ten million views across platforms in its first two days online. But at the time, open-source model capabilities weren't sufficient to complete world simulation, and Tencent's support for the effort gradually diminished.

Later, Che went to Kling AI, then the world's top video model team. The business focused on controllable video generation — making someone walk into frame at the 5-second mark, controlling exactly how they looked and moved. But the problem he himself wanted to solve remained different: can this world run for an hour? If I turn my head, will the person behind me move?

The conclusion from practice: "Even with the era's top-tier video foundation model, you still can't make a world."

A single model cannot simulate the world; world rules and causality require a model plus an engineering system to coordinate — something the gaming industry long ago validated: visuals can be dazzling, but what makes games fun is the mechanism code behind them.

After leaving Kling AI, he joined Huawei's Hong Kong Research Institute AI Lab, responsible for video generation and world model directions. Compared to being an independent researcher at other companies, Huawei allowed him to lead a team and plan technical roadmaps, and gave him his first exposure to thousand-card-scale training.

After joining, to address model quality issues, he reviewed training data entry by entry — thousands of them. By his own standards, 93% were unqualified — shot changes, shaking, excessive motion, or simply near-static images. He proposed a complete solution and worked with colleagues to resolve it.

For this, he was internally granted a team and nearly a thousand cards. It was also there that he first personally designed model architecture and ran pre-training through to completion, on a project called OmniFusion. The key judgment at the time: as the world moves from one state to the next, some changes can be completed at the pixel level, others cannot.

"So modeling must be layered." That World State + Render dual-layer structure grew from this sentence. But what truly shaped XGEN was the roughly 40,000 lines of code he wrote.

In early 2026, he used weekends to build a game called "One Thought, Three Thousand Worlds": a continuously running world, agents living inside it, and corresponding visual presentation.

When he tried conversing with the agents inside, influencing their decisions and lives, he seemed to see the world through others' eyes. In that moment, he felt an intense desire to start a company for the first time.

Because this wasn't a feature that could be neatly packaged into a phone or assistant, nor could it naturally fit into any existing business line. To truly build it required organizing models, systems, products, and teams around this problem from day one.

"This thing can only be done by going out and doing it."

And what big companies taught him wasn't only technical.

At Tencent, he reviewed GameGen-X's ten-thousand-plus videos one by one, conceived the idea, drew the diagrams — yet non-contributors still demanded co-authorship; at Huawei, he received thousand-card resources and team-lead opportunities, but API access, data, manpower, and additional compute all had to be fought for item by item through complex processes.

This is Che's judgment of his own experience, not necessarily the facts as seen by everyone in those organizations. But it shaped how he later built his own organization.

He wanted a more pragmatic, pure environment where he and his team could focus single-mindedly on cracking hard technical problems. After founding XGEN, he allocated far more equity to the team pool than typical tech startups — he calls this the Liu Bang logic: "When wealth gathers, people scatter; when wealth scatters, people gather."

He refers to team members as "experts" and "older brothers and sisters," summarizing his management approach in one phrase: "Holding umbrellas for others."

In 1997, Che Haoxuan was born into a large extended family in Xi'an. His parents divorced when he was in kindergarten: a photographer and a makeup artist, "each pursuing their own artistic lives." He lived with his grandfather, grandmother, and uncle, "forming a very strange kind of living under others' roof." In 2005, his grandmother was diagnosed with lung cancer, and the family's situation began a downward slide; debts for her treatment weren't basically paid off until ten years later.

What he remembers aren't the numbers, but the meat disappearing from dinner tables, and clear-cut rejections.

Around age five or six, he wanted a KFC meal. He spoke up; the answer was clear: no. After that he rarely asked for anything again — "If speaking up won't get it satisfied, then I'll look inward, don't ask, don't think."

He compared his earlier life to the game Temple Run: there's always a monster chasing from behind, and if you slip and fall, it eats you.

With real life cold and sparse, he devoted massive time to novels, anime, and games. From fifth or sixth grade, he read innumerable web novels on a Nokia E71, the buttons and casing worn shiny from constant page-turning.

His favorite was a first-person Diablo fanfiction novel called God's Destruction, 18.65 million words, serialized for over a decade without ending. The protagonist transmigrates into the game, builds a home, adventures, marries and has children. The point isn't completing the game, but continuing to live in another world.

Years later, what Che Haoxuan tries to build isn't a video that plays and ends, but a world where one can keep living.

In XGEN, GEN is generate, and X is the unknown.

Twenty years ago, Che tried to imagine the real lives of people around him, but couldn't continue because he'd never experienced them. Even if XGEN ultimately succeeds, it can only generate one plausible possibility.

But perhaps that's exactly the point of the Door of Anywhere: it can't change the real world someone inhabits, but it can make an unknown world one longs for enterable for the first time.

Twenty years later, Che Haoxuan decided to generate what he couldn't imagine back then.

Cover image: Isaac Levitan, Above Eternal Peace, 1894, Tretyakov Gallery

Powered by @XGEN