So, 24 hours after the "7 x 24" launch, is anyone still talking about Fei-Fei Li?
Fei-Fei Li's World Labs may have been misunderstood.
**
Fei-Fei Li's World Labs may have been misunderstood.
👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon
In 2025, "world models" has become the hottest buzzword in AI, and few names carry as much weight in this space as Fei-Fei Li.
As a professor at Stanford University and a pioneer in computer vision, Fei-Fei Li led the ImageNet project that helped launch the deep learning revolution, earning her the title "Godmother of AI."
Having already achieved enormous success in academia, she chose to start again — founding her own AI company, World Labs.
For the past year, the entire industry has been waiting for World Labs' answer.
On November 13, 2025, Marble — World Labs' first product — was officially launched.

What should have been a straightforward tech release instead became unexpectedly turbulent in the "7 × 24 hours" that followed. From the initial wave of viral excitement, to subsequent skepticism and reflection, to comparisons with the product roadmaps of OpenAI, Google, and Tencent — the episode quickly evolved into a debate about the future direction of AGI.
It all began "7 × 24 hours" ago.
The Sensation Before "7 × 24 Hours"
After its launch, Marble, World Labs' first commercial product, rapidly became the focal point of virtual creation, drawing industry-wide attention.
Put simply, Marble is an innovative tool that allows users to create downloadable, editable, and interactive 3D virtual worlds through text, images, video, and 3D layouts.

Whether for games, virtual reality (VR), visual effects (VFX), or other creative industries, Marble opens up new possibilities for creators. Users can not only build virtual environments but also interact with them, expanding creative freedom.
The product's core concept is "spatial intelligence." This goes beyond AI applications in language or images — it introduces a more groundbreaking idea: a three-dimensional space of "perception-generation-interaction."
In simple terms, Marble enables AI to not just understand and generate flat information, but to perceive and create interactive experiences within 3D space — like a virtual world builder, delivering a new level of immersion.
For example, I uploaded a flat 2D concept drawing of an AI Hacker House:

The world Marble produced did manage to use "as many elements and styles as possible" from the uploaded image to recreate the "world."
However, it's also clear that Marble remains at a relatively early stage, with chaotic and blurry scenes nearly impossible to avoid:

Unlike traditional generation tools, Marble offers multi-layered editing capabilities.
Users can make fine adjustments through natural language instructions — for instance, removing an object from a scene, or making global changes like converting a modern style to retro, or even altering wall structures.
This gives creators more precise control over generated results.
To help users build larger-scale virtual worlds, Marble provides Expand and Combine features. The former lets AI automatically extend the boundaries of existing scenes, while the latter allows users to stitch multiple independently generated 3D worlds into one grand, complex environment.
After creation, users can export content in multiple formats to fit existing workflows, including Gaussian Splats for high-fidelity visual presentation, and mesh formats that can be directly imported into mainstream software like Unity, Unreal Engine, and Blender.

Marble offers commercial pricing plans ranging from a free tier (generating a small number of 3D worlds) to monthly subscriptions ($20, $35, $95), aiming to serve users at every level.

The product launch garnered nearly 3,000 likes and 1.8 million views in a short time, attracting massive attention.

TechCrunch, Fast Company, and The Verge all covered Marble's launch, instantly making it the focal point of the "world model" space and the "top-tier" product representing this new technology.
After the Hype, What Are People Saying?
However, after "7 × 24 hours", browsing through major tech communities and social platforms reveals a clear shift in sentiment: people don't seem to be "buying into" this world model product anymore.
As initial excitement faded, skepticism began to surface.
For example:
"So what exactly is a world model good for?"
"I threw in 10 neighborhood images last night and it still hadn't finished generating by bedtime."

Some offered technical perspectives:
"Tried it, feels like just layered Gaussian splatting."
"Seems like just Diffusion + SfM + Gaussian Splatting."

There were even comments like this:
Andrew Ng, Fei-Fei Li, Yann LeCun — they've pretty much stayed in academia and haven't really accomplished anything in industry.
Initially, World Labs' launch did create considerable waves, attracting large numbers of developers and tech enthusiasts. But as time passed, the discussion seemed to gradually quiet down, replaced by calm reflection and skepticism about the model.
This gap is real.
On communities like X and Reddit, discussion heat around World Labs cooled rapidly. For a highly anticipated product, this "high opening, low closing"舆论 curve forces us to ask:
What kind of product does Fei-Fei Li actually want to build?
The Anxiety of World Models
Marble's birth actually tells a bigger story — one that Fei-Fei Li's team has consistently emphasized: "spatial intelligence."
Three days before the product launch, on November 10, Fei-Fei Li published a new article titled From Words to Worlds: Spatial Intelligence Is the Next Frontier of AI[1], elaborating on the motivations behind the product and once again expressing her "anxiety" about current mainstream large language models (LLMs).

Fei-Fei Li believes that despite LLMs' impressive performance in processing text, writing code, and generating images, they have a clear weakness:
They lack genuine understanding of the physical world.
Existing AI models possess vast knowledge, but they frequently make errors when handling spatial relationships. For example, when estimating distances, judging directions, or imagining what an object looks like after rotation, these top models often perform no better than random guessing.
When generating video, AI also frequently fails to maintain visual coherence because it doesn't understand that there should be a persistent, physically consistent 3D world behind the images.
Simply put, current AI excels at processing abstract symbols but cannot understand physical reality. This limits AI applications in robotics and many other fields.
To address this "anxiety," Fei-Fei Li proposed the concept of "spatial intelligence."
To give AI this capability, World Labs began building "world models" and spatial intelligence — which directly explains why Marble was designed the way it is.
World Labs Is Not Alone
However, "world models" is not a lonely track, and "world model anxiety" is not unique to World Labs.
While World Labs released Marble, it finds itself in a fiercely competitive race, with tech giants including OpenAI, Meta, NVIDIA, Google, and Tencent all sprinting in the same direction.
We've compiled the landscape.
In August 2025, Google DeepMind released Genie 3, explicitly positioning it as a "general-purpose world model." Unlike previous models, Genie 3 emphasizes "real-time interaction."
Where past models typically only generated static video, Genie 3 allows users to enter scenes and interact with them. It supports generating real-time footage at 720p resolution and 24fps, with free navigation — for example, driving or exploring complex terrain from a first-person perspective.

DeepMind even assembled a dedicated "world modeling" team, committed to enabling models to perceive physical environments and simulate changes in water flow, lighting, and more.
Users can not only walk through scenes but also change environments in real time through text instructions — altering weather or adding new characters, for instance. Genie 3's physics-based virtual environment construction is also seen by DeepMind as an important foundation for future training of robots and game agents.
Similarly, Tencent released HunyuanWorld-Voyager in September 2025, an upgrade to Hunyuan World 1.0, and promptly open-sourced it.

Voyager's standout feature is its ability to generate 3D point cloud sequences with depth information from a single static image, combined with user-defined camera paths. This isn't simply flat video — it's scenes with spatial geometric structure.
Tencent also emphasized "long-horizon world exploration" — the ability to maintain stable geometric structure over extended durations or long-distance camera movements, avoiding distortion.


Additionally, it's worth noting that OpenAI's recently released viral video model Sora 2 actually performed "more stunningly" in real-world scenarios.
Despite technical definitions suggesting that Marble and Genie 3's 3D models, which allow users to "walk around" inside them, seem to more intuitively fit the definition of "world models," in popular perception and online buzz, OpenAI's Sora 2 clearly "won more praise" for its presentation of physically realistic effects.
OpenAI itself no longer positions Sora merely as a video generation tool, but emphasizes its stronger "world simulation capabilities."
For example, Sora 2 can simulate the realistic rebound trajectory of a basketball hitting a backboard, attempting to reproduce the deviations that occur in physical reality rather than just generating perfect successful shots.

As a general audiovisual generation system, Sora 2 can create complex soundscapes, human voices, and sound effects with high fidelity.
Moreover, it supports multimodal synchronized generation, with visuals, dialogue, and sound effects aligned.
Beyond this, tech giants like Meta and NVIDIA have also launched their own "world models."
This leaves World Labs in an awkward position: it's doing what may be "the right thing" (building 3D worlds), but other major players are using alternative approaches that seem to be winning users' trust and likes more quickly.
After "7 × 24 Hours," Fei-Fei Li Is Still "Hot"
Now, let's return to our original question:
After "7 × 24 hours," is anyone still talking about Fei-Fei Li?
The answer is: Yes.
But the focus of discussion may no longer be Marble's product hype itself.
We must acknowledge one fact: in today's era of rapid information flow, the short-term buzz around a new product may genuinely only last a brief moment. Technology iterates too quickly — there's always newer tech, cooler products that emerge to replace the previous one.
Those complaints and comments about "too slow to generate" or "this is just Gaussian splatting" are really only directed at Marble as a 1.0 product.
Technology will always be superseded by newer technology.
Marble is just the first product World Labs has put out. It may be imperfect, even technically dismissed as a "frankenstein." But it accomplished something important: it brought the concept of "world models," which had previously remained largely in academic circles, directly to developers and users.
The concept of "world models" itself continues to firmly capture the industry's attention.
So our focus shouldn't be on how many downloads Marble has today, or whether its effects remain merely at the Gaussian splatting level.
What truly deserves our ongoing discussion is the ultimate goal of "world models" that it represents.
Why does this term matter so much?
Because it touches on the most core, fundamental debate in AI: the path to AGI — which way should we go?
The heart of this argument is one question: Should we let AI "read" the world through language, or let AI truly "see" the world?
Path One: Rely on massive data (Scaling).
This is currently the hottest approach, represented by OpenAI and other large language model companies. They believe that as long as we train a sufficiently large model on enough data — all human text, images, and video — intelligence will "emerge."
Sora 2's appearance seems to prove this path, as if it "intuited" some physical rules on its own.
Path Two: Rely on world models.
This is the "academic" path represented by Yann LeCun, Fei-Fei Li, and Yoshua Bengio. They believe true intelligence cannot "spontaneously" emerge from data — it must learn through interaction with the world.

Yoshua Bengio, another "elder" of deep learning and Turing Award winner, also shared his views on "world models" in an interview:
Building world models is one direction for extending LLMs, requiring "world models + reasoning + causality" to compensate for the current limitations of large models that only excel at statistical associations.
Jensen Huang, NVIDIA CEO and an industry representative, also mentioned at a CES industry conference that next-generation AI will move toward "Physical AI" — machines not just processing language and text, but possessing the ability to "perceive, understand, and manipulate the real world."
NVIDIA has also launched "Cosmos," a "world foundation model" platform trained on massive video data to support robotics and industrial scenarios.

They believe AI must build a "simulator" in its "mind" of how the world works. It needs to understand physics (apples fall down), causality (push a cup, it tips over).
Runway co-founder & CEO Cristóbal Valenzuela has also been publicly discussing how "the next step is World Models."
His framing is more direct:
World models are "the next frontier of AI." They won't stop at "predicting the next token" like traditional LLMs, but will seek to understand the physical world for more realistic dynamic scene simulation in film and games.
If AI doesn't have a "body" or "perception," it can never truly understand our world.
However, this path is far more difficult than imagined. Yann LeCun, who has been constantly criticizing LLMs, is now leaving Meta; and Fei-Fei Li, who built ImageNet, has also faced some controversy over World Labs' pace of progress.
Still, Fei-Fei Li's Marble can be seen as an important practice of "Path Two."
It attempts to force AI to understand the world by having it "generate" an interactive 3D world.
This also explains why World Labs is pursuing things that seem even more forward-looking:
[1] Their technical updates for "generating larger, better worlds" aim to solve the scale problem of "world models."
[2] Their revealed RTFM (Real-Time Frame Model) aims to solve the interactivity and real-time performance problems of "world models."

Combined, these elements form Fei-Fei Li's team's complete vision of "world models" — they stand firmly on the side of "AI must understand the world through interaction."
So, after "7 × 24 hours," we no longer need to obsess over Marble's product wins and losses alone.
Fei-Fei Li and World Labs used one product to pull everyone to this most fundamental fork in the road, focusing all eyes on this direction.
The future path to AGI — will it rely on Scaling Law, or on genuine interaction with the world?
Marble's hype may fade, but the debate over "world models" has only just begun. Neither supporters nor skeptics can avoid its significance.
Marble may only be World Labs' starting point, but the idea it represents has already firmly seized the AI field's attention.
So, the next "7 × 24 hours" are what truly matter for the AI field, and what we should really be watching.


References
[1] From Words to Worlds: Spatial Intelligence Is the Next Frontier of AI: https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence