WAIC Observations: Spatial Intelligence in Rough Seas, Everyone Is Looking for the Same Thing
There's a kind of AI that can preserve love.
There's a kind of AI that can preserve love.
👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon
🚦
"The warmth of technology hides in every ordinary story it changes."
I've always been optimistic that technology and AI can help people live better lives.
So I've kept my eye on stories of how technology helps people become more efficient, solves their problems, creates hope and love, and brings new beginnings. Recently, we decided to share these stories at Crossing.
These stories carry power and optimism. They don't just showcase technical breakthroughs and entrepreneurial opportunities — they constantly remind us: even in uncertain times, hope and light still exist.
At this year's WAIC, we discovered a heartwarming story: "A 60-year-old photo studio, and a 70-something photographer."
"Grandpa's Photo Studio Is Open Again"
In the traditional sense, space is simply the physical place where we live. But now, space has been redefined — it has become a digital vessel that holds memory, carrying both personal emotions and collective recollections.
In San Dun Town, Hangzhou, on Chenjiaqiao South Street, there's an old photo studio. Young people are always stopping by to take photos and check in. The walls are covered with yellowed old photographs, like an album that never ends. Every afternoon, seventy-something Grandpa Zhou Quanhu slowly shuffles to the old chair that's kept him company for decades. Sunlight streams through the glass and falls across him as he sips tea and sorts through those precious old negatives.

He always speaks slowly:
"These photos record the smiles of three generations — from the youthful shyness of grandparents, to the sweetness of parents, to the innocence of children today."
But in recent months, Grandpa Zhou's routine has changed. He's developed lung disease and can't visit the studio as often. Grandpa Zhou often says the photo studio is a place where many people come to find beautiful memories, and he's reluctant to close it down easily. So he decided to hold a "living farewell" — right there in that little shop full of memories.
A camera can record a moment in time and space, yet it seems no one can record this photo studio itself. This is Grandpa Zhou's only regret. From when he started as an apprentice at 25, this studio has been quietly running for decades.

Right then, an engineer named A Hang heard about Grandpa Zhou's story.
A Hang works at Manycore Tech, a spatial intelligence company. He rushed over from the other side of Hangzhou, felt the light and shadow of this photo studio, and understood that every photograph here is a fragment of memory from a warm time.
Using his company's research on "3D Gaussian technology" — an AI technique that can completely replicate real spaces inside a computer — he made a "digital backup" of this old photo studio.
Grandpa Zhou's Digital Photo Studio: https://www.kujiale.com/pub/koolab/koorender/gifts

Generated by 3D Gaussian technology. Simply put, it means creating a copy of this shop in the cyber world, so people can still "walk into" it on their computers, see those old photos, and feel the warmth of those years. This isn't just about preserving a small shop — it's more like building an eternal home for a community's warm memories.
Here you can also feel the earnest hopes of netizens: "Of course it exists, and I believe it will keep being passed down."

Grandpa Zhou's "Digital Photo Studio"
When Grandpa Zhou calmly said, "Either you forget others, or you're forgotten by others," the AI engineer used technology to write a third possibility:
Let beautiful things never truly disappear.
I think the most beautiful meaning of technology is to serve love.
Beyond its nostalgic use, 3D Gaussian also represents a technical vision of "copying the real world 1:1 into the digital world."
When you "walk into" this digital photo studio, you'll find everything identical to the real thing: you can see the details of every photo on the wall, feel the warmth of sunlight streaming through the glass.
This is like building a bridge between the physical and digital worlds.
This bridge has a name:
🚦
Spatial Intelligence
It allows machines to not only "see" the world, but to "understand" space — to grasp the stories and emotions hidden in every corner.

What Is Spatial Intelligence?
Simply put, "spatial intelligence" means enabling AI and robots to understand and operate in the 3D world we live in, just like humans do.
Let me give you a real-world example.
Many of you have robot vacuums at home. Unlike humans, who can find their bed with eyes closed and know to go around table corners without hitting their legs, these seemingly natural abilities are actually huge challenges for robots.
They often run into your pet's 💩 in the bathroom, completely oblivious, then drag it all over the house...
Robots can see obstacles but can't understand what objects mean or how they relate to each other. They don't know what can be touched, what can't, what's dangerous, what's precious.
Spatial intelligence is about truly teaching them to "live" in a 3D world — like teaching a child to understand the world — so AI can not only see tables, chairs, and walls, but understand their purposes, their spatial relationships, and how to interact with them safely and gracefully.
The Spatial Intelligence Industry: "Big Winds, Big Waves, Big Fish"
Some companies are pursuing frontier breakthroughs in spatial intelligence. Others are quietly building the infrastructure for this field.
At this historical inflection point where "AI moves from the digital world to the physical world," every company is writing its own footnote, all believing in a simple business truth: "The bigger the storm, the more expensive the fish."
The more challenging an emerging market, the more likely it is to birth billions in value.
1) Anyone Could Catch the "Big Fish" at the Crest
Spatial Intelligence
Fei-Fei Li — renowned AI expert, Stanford professor, "Godmother of AI" — with so many titles already, added another last year: founder of WorldLabs. Starting from her ImageNet research with over 80,000 citations, she has been obsessed with making machine systems "see" and "understand" the real world.
When Fei-Fei Li founded WorldLabs, it immediately became the most closely watched spatial intelligence AI project. She put it directly: "Spatial intelligence will be the next frontier of AI."
This startup's goal is: to build Large World Models (LWMs) that can perceive, generate, and interact with three-dimensional worlds.
They released an early result of their AI system that generates 3D worlds from a single image late last year:
Within just a few months, WorldLabs completed roughly $230 million in funding at a valuation exceeding $1 billion. Lead investors include Radical Ventures, a16z, NEA, NVIDIA's venture arm, and even Google's AI luminary Jeff Dean and Geoffrey Hinton — winner of the 2024 Nobel Prize in Physics and the 2018 Turing Award — expressed their support.
So why is spatial intelligence attracting such attention? In both academia and industry?
The reason is simple: whoever masters spatial intelligence first could seize the commanding heights of next-generation AI applications. The scenarios it can unlock are simply too numerous.
Fei-Fei Li has also spoken about spatial intelligence's application prospects:
Just as written language transformed human communication, spatial language will change how AI interacts with the physical world. From content design, architecture, industry, art, gaming to the metaverse — everything could be transformed. I believe the fusion of hardware and software is coming, and that's a very promising application scenario.
But this is just the tip of the iceberg.
In embodied intelligence, spatial intelligence is core among cores. With it, robots can truly "understand" the world and know how to create productivity in real environments.
Embodied Intelligence
Spatial intelligence is the foundation of embodied intelligence. Without spatial understanding, there can be no operation in the physical world. Humanoid robots are the physical carriers of embodied intelligence.
You can understand their relationship like this: spatial intelligence lets robots know how tall a table is, how to navigate around obstacles, and how the same object looks from different angles.
With this "common sense," robots can move through the real world as naturally as humans.
Compared to the relatively concentrated spatial intelligence track, the embodied intelligence field is "a hundred flowers blooming," appearing more "vibrant."
Numerous familiar players have emerged at home and abroad, each with their own unique approach and application scenarios.
Crossing recently surveyed China's top embodied intelligence players: "After Allen Zhu's Wake-Up Call, We Surveyed the Survival Status of 10 Leading Humanoid Robot Companies."
Of course, beyond these companies we've already covered, many more "vibrant" enterprises await our discovery.
These days, the most buzzworthy, hottest event in embodied intelligence is naturally Elon Musk's Tesla restaurant opening, with Optimus robots "selling popcorn at the stall."
Those dexterous hands that can sell popcorn and perform precise operations — behind them lies spatial intelligence technology.
World Models
In practical embodied intelligence applications, we can make a "not entirely apt" analogy: spatial intelligence and world models work together like left and right hands.
Spatial intelligence answers "where am I, what obstacles surround me," while world models predict "if I do this, what will happen."
When you prepare to stand up from a chair, your brain automatically rehearses the motion: how to position your feet, how to shift your center of gravity, whether to use the table for support. This "rehearsal" ability is what world models do.
Combined, AI can autonomously act in 3D worlds just like humans.
This is why top researchers like Fei-Fei Li and Yann LeCun — one of deep learning's "three giants" — both believe that language and text alone cannot build true intelligence. A model that transcends both flat images and language, one that truly captures three-dimensional structure and understands spatial intelligence — a world model — is the core development direction for next-generation AI.
"Understanding the 3D world, generating the 3D world, learning to reason and act in the 3D world — this is the fundamental proposition of AI." — Fei-Fei Li
Just last month, the "last holdout" for world models, Yann LeCun, released new progress — the PEVA model, giving embodied intelligence human-like "anticipatory ability," though it only achieved 16 seconds of coherent scene prediction so far.

2) The "Waves" Are Big, and Fierce
Anyone in spatial intelligence could catch the big fish at the crest, but how big the "waves" are, and whether you can withstand them, tests the resolve of every individual and every startup.
This is a marathon of thousands, where everyone started running before even hearing the starting gun — and immediately ran into some obstacles.
Where is the data spatial intelligence needs?
Even for WorldLabs, founded by a top scientist like Fei-Fei Li, which has already achieved initial commercial success and is surrounded by "possibly the best scholars in the world," the results they put out still get "nitpicked" by some observers:
Let's be honest, what we can see now is basically something that looks like a "3D game."
Simply put, after experiencing the "bull rush" of large language models in recent years, people are a bit "dissatisfied" with the slow pace of 3D AI research.
This actually points to the biggest problem behind spatial intelligence: 3D spatial data is hard to obtain.
Fei-Fei Li has repeatedly discussed the same problem in various interviews:
One of the biggest challenges in spatial intelligence is: there's abundant language data on the internet — but where is the data spatial intelligence needs?

Robot "Clumsiness" Is Surprising
From theory back to reality, let's look at what embodied intelligent robots are facing without sufficient 3D data.
Contrary to the grand theoretical descriptions, current robots have some rather "surprising" and "embarrassing" problems: they can't fold blankets.
This is a fascinating example — the "blanket-folding dilemma," raised by Manycore Tech co-founder Xiaohuang Huang at last year's GeekPark conference.
Here's the situation: when you tell a robot "help me fold the blanket," it can indeed use its camera to distinguish which blanket is folded and which is messy. But it simply doesn't know how to actually do it.
Moreover, even if it learns to fold one blanket, give it a different shape and it might be lost again.
What does this show? Embodied intelligence needs massive amounts of "practice," and the difficulty of this training is extremely high.
Whatever the Training Method, Massive Data Is Needed
Currently one of the most efficient ways to train embodied intelligence is Sim2Real (simulation to reality).
Simply put, it's about letting embodied intelligence achieve enlightenment in a "Longchang" — the virtual environment — accelerating training by hundreds or thousands of times on "mappings of real-world scenarios in virtual scenarios."
NVIDIA, for instance, specifically developed the Isaac Sim simulation platform to provide robot simulation environments.

While Sim2Real methods are extremely efficient for training in simulation, real-world physical scene data remains essential.
Everyone Is Looking for the Same Thing
You could say the robot's brain is in the digital world, but its body is in the physical world.
To solve this kind of problem, the key is establishing a "data bridge" between the physical and digital worlds.
Looking back at those top scientists, startups, and well-funded tech giants mentioned earlier — they all face the same problem:
Where do we find this data bridge?
This is the core dilemma currently facing the entire AI robotics industry — with the best algorithms in the world, without enough data to "feed" them, robots still struggle to truly understand the three-dimensional world we live in.
The "Bridge" Between Digital and Physical Worlds
When training large language models, the need for more, better, "less toxic," cleaner data gave rise to "a generation of Scale AIs."
In the spatial intelligence era, we're hoping to see a similar generation of companies "blooming across the flower fields."
The primary challenge currently facing embodied intelligence is the lack of real-world physical data, especially physical properties like mechanics and spatial relationships. When AI-equipped robots enter the physical world, previous one-dimensional and two-dimensional data become severely limiting.
For AI robots to understand the complexity of the three-dimensional world, they first need to learn three-dimensional world knowledge — what's often called "3D interactive data."
We noticed at WAIC that one domestic embodied intelligence player has particular advantages in this area — Manycore Tech, the team behind the 3D Gaussian technology for Grandpa's photo studio. More interestingly, they're also a team we frequently "encounter" when researching the spatial intelligence and embodied intelligence industries.
I say "encounter" because when you study technological developments in this field, you always see them.
In March this year, after Manycore Tech's team released their spatial model SpatialLM, it immediately climbed to the top three on the open-source community Huggingface's trending list.

Their dataset was also acknowledged by name in a joint paper by Google DeepMind and Stanford.

Spatial Intelligence Cannot Be Created from Nothing
Actually, as early as 2018, Manycore opened its 3D scene dataset InteriorNet to the tech community — one of the largest indoor spatial deep learning datasets in the world at the time, and still today.
As we understand it, this dataset is called InteriorNet because it serves as "the 3D version of ImageNet," paying homage to Fei-Fei Li, the mother of spatial intelligence.
This shows they invested early in embodied intelligence, a field of "big winds, big waves, big fish."
After this, Manycore's spatial intelligence platform SpatialVerse has been continuously exploring how to make synthetic data carry more realistic and richer physical information. For example, during this year's WAIC, Manycore's newly released 3D Gaussian semantic dataset represents one such exploration — containing 1,000 3D Gaussian semantic scenes, each with structured information on spatial relationships, positions, and object dimensions.
During WAIC, Crossing had the opportunity to conduct an on-site interview with Chief Scientist Rui Tang.
👦🏻 Koji
As Manycore Tech's chief scientist, how do you view AI entering the real world? What changes will it bring?
👦🏻 Rui Tang
I believe AI will definitely enter the physical world in the future. Internally, we actually categorize AI into "screen space," which is the "digital world," and beyond that, the "physical world."
Currently, major companies like DeepSeek, Kuaishou, and Tencent are still mainly focused on the data world — for instance, completing a service with a single voice command. But we believe what's more important is pushing AI from the data world into real physical spaces. For example, from factory workers and office assistants to home services, AI has enormous potential in these areas.
👦🏻 Koji
Manycore Tech has extensive practice in synthetic data. How do you evaluate its value and limitations in AI model training?
👦🏻 Rui Tang
The advantages of synthetic data are obvious.
Real-world data collection is expensive, requiring significant hardware investment and human coordination. In virtual environments, we can construct thousands of consistent scenarios — automatically controlling object positions, lighting changes, materials, and so on through programs. Through this method, we can achieve high degrees of randomization and scale effects.
Of course, it definitely has limitations — for example, it can't completely 1:1 replicate the distribution characteristics of the real world. We're also working hard to break through this problem, building bridges connecting digital and physical worlds, using spatial computing capabilities to bring AI into real environments.
👦🏻 Koji
What can Manycore do as AI moves toward the physical world? What unique value will it bring?
👦🏻 Rui Tang
Manycore actually began exploring spatial computing back in 2011. At that time, our founder had just returned from NVIDIA, where he had been working on parallel computing research.
At this point, we faced a key choice: use parallel computing to simulate the digital world, or the physical world?
AI technology wasn't mature enough then, so we chose to use parallel computing to simulate the physical world — which became the starting point for our later spatial design-related technologies. After developing the tool "Kujiale," we attracted a large user base, and in 2016, we further recognized the potential of this spatial data.
So we developed InteriorNet, with 16 million sets of pixel-level labeled data, plus 15,000 sets of video data — roughly 130 million images total — and released this dataset as a tribute to Fei-Fei Li's team's ImageNet.
People have long said there are "three carriages" in AI: computing power, algorithms, and data.
Among these, we believe "data" is the most core resource — like oil, like electricity — capable of driving algorithm optimization and evolution. Our strategy is to first use tools to generate spatial data, then use this data to drive spatial intelligence algorithm R&D, and finally use algorithms to feed back into the tools. This forms a closed loop. You could say this is also Manycore's unique position in the spatial intelligence track.
👦🏻 Koji
Manycore's position in the ecosystem is both unique and important, truly enabling AI to "understand the three-dimensional world." Thank you, Professor Tang.
👦🏻 Rui Tang
Thank you.

Manycore Tech's digital photo studio recreated with 3D Gaussian technology
From this conversation with Manycore Tech Chief Scientist Rui Tang, we learned how they've transformed years of accumulation into a unique business flywheel:
[1] Spatial design tools (Kujiale) generate massive scenario data, accumulating indoor spatial datasets;
[2] → This large data foundation enables training spatial models;
[3] → Models feed back into tool optimization, further exploring spatial intelligence;
[4] → Forming a closed loop.
This loop rolls like a snowball, growing bigger and bigger, ultimately giving Manycore Tech some hard-to-replicate advantages in spatial intelligence.
This reminds us of other similar cases:
Just as Kuaishou accumulated massive short-video assets, ultimately enabling video generation AI like Keling AI; miHoYo's years of game production experience gave birth to Honkai: Star Rail's CG.
You could say what Manycore possesses is precisely what the AI industry lacks most: deep mining of existing business, rather than creation from nothing — solving the problem of "difficulty in landing AI in the physical world."
The Things That Never Disappear
When we reopen that digital photo studio link, we discover something interesting: that little shop is still there waiting for every visitor.
The photos on the wall are still those photos. Sunlight still streams through the glass windows. Even the air still seems to carry that old camera smell. Only now, this space is no longer bound by time — it can be "visited" by countless people simultaneously, its doors open to anyone at any time.
When we can reproduce the world at millimeter-level precision, the boundary between real and virtual begins to dissolve.
This finally gives us a new way to preserve memory, connect emotions, and understand the world.
🚦
Embodied intelligence is completing AI's body; spatial intelligence lets it truly stand in the three-dimensional world and see.
Grandpa's photo studio is where these technologies become soft, intimate, and tangible.
It's no longer just a photo studio.
It's a starting point born from 3D Gaussian, an entrance where we begin rebuilding the bridge between digital and physical.
It never closes.
Everyone can push open the door, sit down, and leave a moment of their life there.

