10 Q&As: A Deep Dive into the Divides in Embodied Robotics — A Conversation with the Founders and Chief Scientists of Sudu Tech, Ant Group's Lingbo, Independent Variable Robotics, and Poke Robotics
Where does the data come from? How is the model built? Where does it get deployed?
Where Does the Data Come From? How Do You Build the Model? Where Does It Land?

👦🏻 Host: Koji
🥷 Editing: Crossing
🧑🎨 Layout: NCon

This is a rare public discussion — one without convergence — about the route-level disagreements in embodied AI.
In 2026, embodied AI isn't short on hype. What's missing is someone putting the disagreements on the table. Where does the data come from? How do you build the model? Where does it land?
If OpenAI and Anthropic move into embodied AI, would that be a dimensional-strike takedown? Where are the limits of Astra's capabilities? Is "building an airplane as easily as building an app" a fantasy, or within reach?
On the morning of September 10, at the 2026 Inclusion · Bund Summit, a panel titled "Physical Intelligence — Divergence and Decisions" took place. The panel was hosted by Koji, with Zheng Han, co-founder and CEO of Sudu Tech; Yujun Shen, Chief Scientist of Ant Lingbo; Qian Wang, founder and CEO of Independent Variable Robotics; and Huazhe Xu, founder of Poke Robotics and assistant professor at the Institute for Interdisciplinary Information Sciences, Tsinghua University.
2026 Inclusion · Bund Summit | "Physical Intelligence — Divergence and Decisions" panel
Through this conversation, you'll learn: 1) the core assumptions and risks behind different technical routes; 2) how to tell which metrics are leading indicators; 3) the difficulty language model companies face entering embodied AI, and where startups' moats lie; 4) the short-term vs. long-term logic of embodied AI commercialization; 5) which areas are overhyped or undervalued — potentially corresponding to investment opportunities or risks.
When an industry stops arguing about "what's right" and starts proving "what works" through products and deployment, that's when a sector truly moves from bubble to industry.
Listen on WeChat:
Listen on Xiaoyuzhou:

🎬 The video is also available on @Koji杨远骋's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms:
Below is the transcript, edited for clarity.
The Data Debate: Simulation or Real Machines?
👦🏻 Koji (Crossing)
Let's get straight to it. Today we'll cover three topics: the data debate, the intelligence debate, and the deployment debate.
Why all these "debates"? Because in 2026, we're standing at a crossroads — there are many paths that haven't converged, waiting for people to explore.
First question for Zheng Han: on data, some people bet on simulation, others firmly bet on real machines. You firmly bet on simulation — what signals did you see back then? And what new reinforcing signals have appeared today?

Koji, Founder of Crossing
👨🏻💻 Zheng Han (Sudu Tech)
We firmly believe in using simulation. But at the same time, we also firmly believe real-world data is extremely valuable.
Our starting point is pretty simple. We began building robot simulators and synthetic data back in 2015–2016, and then founded a company to commercialize them — we've been through two rounds of iteration.
First, research and commercialization are different things. In research, you can experiment with open-source projects, data, and algorithms. But typically, after people try it out, they don't know how to scale up and modify the technical models — so it just stops there.
The reality is, whether it's model-based reinforcement learning or large-scale pretraining — whether it's the simulators we previously open-sourced, CPM and Skill, or NVIDIA's Isaac — from our internal perspective, they all have certain limitations.
So after 2024, we decisively built a large-scale, fully re-architected closed-source system, including the data pipeline and simulator pipeline — simulation, including model-based RL and large-scale reinforcement learning inside the simulator.
Zheng Han | Co-founder & CEO of Sudu Tech
In our April demo this year, we showed that for certain skill operations under strong generalization across combinations of objects, environments, and lighting — tasks like grasping, placing, and assembly — we can achieve nearly 100% success rates with almost no restrictions on objects or environments.
On data, we insist: simulation is unavoidable. Because relying purely on real-world data collection is slow, and the quality is sometimes limited.
Recently, some peers in the United States have also started strengthening their simulator capabilities. Their methods differ from ours, but it's a signal to the industry.
👦🏻 Koji (Crossing)
Lingbo has clearly chosen real data. Yujun, in your view, which aspects of real data can't be replaced by simulation? What's the biggest value it provides compared to simulated data?
👮🏻♂️ Yujun Shen (Ant Lingbo)
I wouldn't say we've firmly chosen real data. We've never felt that one type of data must be chosen or must be excluded.
Different data plays different roles at different stages. For example, in the pretraining stage, we think internet data and data collected from the real physical world probably matter more. But when it comes to a specific task, simulated data still has value — autonomous driving is a case in point.
Yujun Shen | Chief Scientist of Ant Lingbo
Why do we insist that part of pretraining data come from real scenarios?
Imagine one day a robot actually goes to work in the real world. The way it works will be different from the digital world. Its only observations must come from the sensors on its body — we can't plug a keyboard into a robot and type input into it. And real-world sensors inevitably have errors. There's no perfect sensor — that's a physical constraint. In that situation, the robot can only rely on its onboard sensors as input, and then act.
Simulation may have a problem now — people talk a lot about the Sim-to-Real gap, but from first principles, I think the bigger issue is actually the scalability of Real-to-Sim.
In other words, can we really bring sensor noise and other lower-level things into the simulator faithfully — and can that process be replicated at scale? If Real-to-Sim can't scale, then just talking about the Sim-to-Real gap gradually shrinking may not be worth much.
As for which data must come from the real world and can't be replaced by simulation — at a deeper level, it's because robots must work in the physical world, and many imperfections of the physical world are hard to simulate.
👦🏻 Koji (Crossing)
Give a concrete example — what kind of imperfection is really hard to reproduce in a simulator?
👮🏻♂️ Yujun Shen (Ant Lingbo)
Tactile information, tactile signals, for example.
Many tactile sensor companies now offer tactile simulation, but we've tested it — the signals from real sensors are wildly different from the simulated ones.
Signal frequency, amplitude, consistency across sensors — these are very hard to replicate perfectly inside a simulator.
The Intelligence Debate: Will Astra Crush Embodied AI?
👦🏻 Koji (Crossing)
After GPT-6's Astra launched recently, social media has been flooded with tests of people using Astra to control robotic arms on all kinds of tasks. There's even a claim that embodied AI is about to be crushed by general large language models — a dimensional strike. Qian Wang, how did you feel watching those videos?
💂♂️ Qian Wang (Independent Variable Robotics)
This is indeed the hottest topic in the industry. Just this morning I saw another video of someone using Astra to control a dexterous hand.
This version of the model actually had a lot of robot data trained into it. Since early last year, OpenAI and Anthropic have been purchasing robot data in large quantities, and they run small-scale data factories of their own to collect robot data. So it's not particularly surprising — it was expected.
Qian Wang | Founder & CEO of Independent Variable Robotics
I've always had a "hot take," and colleagues on the front lines of large models would probably agree:
Today, the substance of intelligence lives in the data — well, partly in the evaluation too. The model is more of a distiller and a container: during training it's a distiller; during deployment it's a container.
When we build a frontier AI system today, the model's role has actually become extremely small. What matters more is the data — you need better and better data. How you create data, select data, validate data, label data — it's precisely these things that drive AI systems forward.
So the question becomes: if OpenAI and Anthropic want to do embodied AI, do they need the same data infrastructure, validation infrastructure, and other infra that we do? Do they need to collect data from the real world, build simulators, run evaluations in the real world, integrate with hardware, and so on?
If they still have to do all these things, then they're just another embodied AI company. Because these difficulties don't diminish with their advantage in foundation models — in fact, embodied AI companies may have a substantive advantage.
👦🏻 Koji (Crossing)
So you believe GPT-6 Astra won't deliver a dimensional strike on the embodied AI industry?
💂♂️ Qian Wang (Independent Variable Robotics)
The real worry people have is: can large language models generalize their universal intelligence from other domains to embodied AI, and deliver a dimensional strike?
I think it's unlikely.
First, nothing like that has ever happened in the history of language models. Same with 3D recently, and computer use — even though video generation models are excellent today, that hasn't directly translated into an inherent advantage at the action level, at the robot control level.
Conversely, the lesson for us is: pretraining really does work.
Many people in the industry worry: we've done massive pretraining and added lots of general data — is that actually harmful? Does it hurt model performance? With Astra, we see it's at least a bonus — at least not a bad thing. We should firmly keep going down this path.
Finally, the only gap between us and Anthropic or OpenAI may be this: they've integrated this infra into a general-purpose large model, while we're building specialized models.
The reality is, a general language model can't run inference and deploy on a robot at a reasonable speed. We genuinely need a dedicated embodied model to do this — it's the more efficient approach.
👦🏻 Koji (Crossing)
Poke Robotics recently released a video of a robot making mapo tofu. Huazhe, in your view, could Astra complete a task like that too? More fundamentally — does a robot completing a specific task prove that its underlying model has truly gotten stronger? If not, what metrics should we look at?
🤵🏻♂️ Huazhe Xu (Poke Robotics)
Astra is genuinely impressive. We tested it internally the day it became available.

Huazhe Xu | Founder of Poke Robotics; Assistant Professor, Institute for Interdisciplinary Information Sciences, Tsinghua University
If you think of embodied tasks as requiring semantic generalization, spatial generalization, plus physical generalization — Astra is relatively weak only on the physics side; it's relatively strong on semantics and space.
We give Astra a prompt and it recognizes everything on the desk — any category — and it always moves to the right place. With previous VLM or large models, the robot would basically get there, start flailing around, and that was it.
But one unique thing about Astra is that it can actually grasp — and it can grasp very hard-to-grasp objects. For example, a chopstick lying on a table — for a robot, such a small object usually has to be pushed against the table edge, maybe sliding a bit, before it can be picked up. Astra can pick it up now.
At that point we wondered — maybe rotation is its weakness? So we tried having it stand the chopstick upright and stick it into a cup. It could do that too. So it's not just two-dimensional space — it performs well across the whole three-dimensional space. But when we asked it to do complex tasks involving physical contact — like prying something open with chopsticks — it fell short.
I think this is the last stronghold of embodied robotics companies: understanding physical change.
We also used Astra to try stacking boxes and folding clothes — things embodied companies love to test — and Astra couldn't do any of them. And among all the extremely complex tasks, making mapo tofu is surely one of the hardest.
Back to Astra: what it can currently do is relatively spatial-semantic embodied tasks. Among complex contact-rich tasks, the things it can do are: putting a chopstick into a cup, and handover — picking something up, passing it to the other hand, then placing it on the other side. We tested that many times; it succeeded once or twice.
What Astra made me realize is — a bit like what Professor Mengdi Wang (Director of the AI Innovation Center at Princeton University) said in her keynote earlier — embodied models should also be connected to a Codex-like system, or plugged into a harness environment.
Because embodied models work in general scenarios, but the harness needs strong semantic capability. The stronger the harness's semantics, the smarter the embodied model can be used.
Take grasping fruit. When people train models, they usually start with foam fruit — but that data is very harmful: foam fruit can be pinched up with a bit of extra force, digging in slightly; but you absolutely cannot grab real fruit that way. So we want the model to check first: if it's foam, use the old grasp; if it's real fruit — say a pineapple — find a way to grab the green leaves on top instead of gripping the body.
Most embodied models don't have this recognition ability. We need the model to first use CoT (Chain-of-Thought) to tell the robot whether it's a real pineapple or a fake one, and only then grasp.
So I believe the endgame should be a complete system where the large model, the embodied model, and the physical world are connected together.
👦🏻 Koji (Crossing)
Yujun, in your view, is the capability Astra is now demonstrating some kind of signal of a scaling law for embodied AI today?
👮🏻♂️ Yujun Shen (Ant Lingbo)
My view is close to Huazhe's, maybe a bit more radical. I think embodied AI will gradually influence the development path of large models and cloud models.
First, the way embodied models ultimately work must be deployed on robots — which is different from the "turn-based dialogue" of digital-world models.
In the digital world, you ask it a question and it answers — while it's answering, you can't keep interrupting. The interaction is linear. But robots act based on continuous input from their onboard sensors — the input streams in nonstop, even while the robot is mid-operation, and the robot needs to adjust its reasoning at any moment based on new input.
I believe embodied models may end up becoming tools used by higher-order models.
As a tool, it needs its own training paradigm — because it must handle a scenario of continuous input and continuously output a policy.
Why do I say embodied scenarios might offer new development directions for higher-level, interactive expert-type models?
Here's an example: human-to-human conversation is actually multi-modal — not just language and vision, but many external factors. Say I'm chatting with someone and it suddenly starts raining outside. Seeing the rain might change my mood, change how I talk. Or I notice the person looks unwell, so I swallow half the things I was about to say and switch to something else.
Today's cloud-based digital-world models haven't considered this, because they don't need to yet — the interaction happens behind a screen; everyone just communicates through vision and language.
But in the future, if robots really converse with people, the amount of information they can take in will be enormous — and much of it will come from real-world sensors. We'll observe the weather, sense the temperature... all of this will affect the model's output.
In the future, as robots gradually enter daily life and bring back more data about the physical world and real life, this data will become new corpus for large language model companies, making them smarter.
Robots will no longer be cold agents hiding behind a screen — they'll be agents that, like humans, can feel their surroundings and adjust how they talk to people based on changes.
👦🏻 Koji (Crossing)
The three of you are broadly optimistic — seeing opportunities and inspiration everywhere. Zheng Han, anything to add?
👨🏻💻 Zheng Han (Sudu Tech)
It's actually converging with a point I've been making: things will stratify top and bottom. The upper layer of abstract understanding will stratify.
I've been saying this for years: Facebook AI Research (FAIR)'s biggest rival is actually OpenAI — and recently Anthropic too. Another DeepMind robotics researcher also just joined Anthropic.
One more thing: the foundation of robotics is skills. The generalization and high success rates of low-level skills must match the level of the upper-layer model. You can't have low success rates with some generalization — that's a fatal problem when connecting the upper and lower layers.
So when embodied companies interface with upper-layer models, they must watch out for this: generalization at high success rates.
Late last year, in our demo, we did language interaction with the Qwen model — it understood the surrounding environment and objects, and finally broke things down into operational steps generated through reasoning. But it absolutely depends on extremely reliable, stable low-level manipulation of the environment and objects to complete the full training of the action.
This year the idea has come back to the foreground — I think that's the big trend, especially after Astra appeared.
The Deployment Debate: What Metric Proves Embodied AI Has Moved Past the Bubble?
👦🏻 Koji (Crossing)
There's some real-world skepticism about the embodied AI industry, mostly about commercialization: lots of demos, hot fundraising — but where are the scenarios that can truly be replicated at scale and keep customers paying?
This question is for the three CEOs of embodied companies: if you could only look at one metric, which one best proves embodied AI has moved past bubble skepticism and truly entered the industrial stage?
👨🏻💻 Zheng Han (Sudu Tech)
What research and commercialization demand of robots is exactly the same as what industrial robots and automation demanded before — high success rates are the prerequisite for the vast majority of tasks.
In commercial scenarios, most tasks require robot success rates close to 100%; only a very few can tolerate an 80% success rate or allow second-attempt corrections.
On that premise, what distinguishes us from automation equipment is: we also need a degree of generalization across objects and environments — and we can't compensate with massive post-training. Otherwise, every new scenario requires heavy overfitting, and that defeats the whole point of embodied AI.
So when it comes to commercialization, our biggest test is first in the algorithms: generalization at high success rates. And of course, the reliability of the hardware and how you deliver service — those challenges can't be avoided either.
💂♂️ Qian Wang (Independent Variable Robotics)
If I have to name one metric, it's this: customers sustainably paying — paying for frontier technology.
That's the ultimate criterion, and the most realistic one.
A lot of today's commercialization isn't actually about customers sustainably paying. It might be a collaboration of a specific form at a specific historical moment. Or it might carry over from a previous era — traditional automation, the last generation of robots.
But people should really think about: what is this generation of robots better at than the last? What real value can it deliver to customers?
This is already starting to happen. Recently, in logistics scenarios, we've genuinely done things traditional robots couldn't do — and exceeded human ROI levels. This kind of thing — truly deployable, sustainably generating productivity — is what matters most.
On the other hand, a company must make money. With money, you can build bigger models and keep exploring the AI frontier.
Language model companies went through a similar path. At first they made no money at all; then at a certain stage, when they became genuinely useful, they began actively commercializing. Commercialization and pushing the frontier are mutually reinforcing processes.
In a sense, commercializing too early is harmful. But we may have passed that window — we're now entering a phase of flywheel effects, where the two reinforce each other.
🤵🏻♂️ Huazhe Xu (Poke Robotics)
Embodied commercialization comes in several categories.
One is post-training: I partner with a certain brand or factory and run a closed loop inside it. For example, I set up an oden snack stand outside my home — the core metrics are how many hours it takes me to deploy it stably, and how many snack stands I can open. That's the same deployment logic as traditional robotics — it's about efficiency and ROI conversion.
Second, I think this wave of embodied AI has another commercialization metric, one similar to the commercialization path of large language models. In the short term, there may be no visible economic returns — but it will be a "step-function" commercialization path.
Why does everyone feel, viscerally, that large language models are useful?
Because the cost of tolerance for errors is different. If GPT outputs an answer with a 50% success rate, I just chat with it a few more rounds and I'll most likely get a good result. But if a robot completes a task with a 50% success rate, that means there's a 50% chance it smashes a cup — and everyone will think it's useless.
Even so, we must unswervingly pursue the large-model-ification of embodied AI, because that current 50% success rate is 50% across all tasks. We wait for it to climb from 50% to 60%, 70%, all the way to 99.9% — and the commercialization at that point is the true commercialization of this wave of embodied AI.
That's what we're really looking forward to: the foundation model reaching a certain stage in a single leap, and starting to devour huge parts of the physical world.
Looking Back from Five Years Out: What's Overhyped, What's Undervalued?
👦🏻 Koji (Crossing)
Last question. There's a saying: people always overestimate the change technology will bring in the next two years, but underestimate the change in the next ten. If we look back at this moment from five years in the future, what do you think is undervalued in embodied AI today? And what's overhyped?
👨🏻💻 Zheng Han (Sudu Tech)
I think the most overhyped thing is: the sophistication of China's supply chain.
There are certainly many options, but for the supply chain embodied AI now needs — one mature enough to support our actual iteration speed — there's still some work to do.
First, we need some volume to support it. Second, many sensors and key components in the supply chain are scattered across mature industries — the automotive supply chain, the industrial robot supply chain. When it comes to embodied AI specifically, the current supply chain still lags behind large-scale industrial manufacturing. Though it can definitely be fixed.
As for what's undervalued — responding to what Huazhe said earlier about reaching 99.9% success rates taking a long time — the difficulty is something people shouldn't be too pessimistic about.
I think in the next 1–2 years, people will see entire categories of tasks in many industries being handled by robots.
Of course, if you mean robots handling 99.9% of all tasks — that will indeed take a long time.
👮🏻♂️ Yujun Shen (Ant Lingbo)
Embodied AI is a complex systems engineering problem. People may still be underestimating the system's capabilities, while tying many problems — whether it can be deployed, the cost of solving them — to model performance. But that's not right.
Because commercial robotics ultimately comes down to cost, the robustness of the whole system, the business model, and so on. The model is just one link in the chain.
Take deployment: merchants will definitely look at cost. Do I really need this particular robot? Do I need a full robot? Would just a pair of arms do?
As for what's overhyped — I think we're currently overestimating "the technique of training models."
I've heard voices saying embodied AI is about to be crushed by digital-world models — as if you could just gather a batch of physical-world data and train an embodied model using the old training methods. I don't think that holds.
Embodied AI may follow the same pattern as the digital world, but it definitely won't follow the same technical route. Because the final application scenario determines how the model differs. The digital world is conversational, but embodied AI must reason while the robot works.
Embodied AI will have its own new training paradigm — one that supports continuous processing of physical-world information, maybe 24/7, even 365 days a year, nonstop.
💂♂️ Qian Wang (Independent Variable Robotics)
Looking back from five years out, I think people will feel we still underestimated the importance and potential of embodied AI.
Even though today we already recognize embodied AI as important — a very important component of AGI — I don't think people truly understand it yet.
Following on the Astra question — why does data matter?
Because it genuinely contains intelligence that can't be obtained any other way. Complex physical interactions, all kinds of physical processes, and as Shen mentioned, sensor noise and the inherently uncontrollable aspects of the physical world — you really can't get these from video or 3D, let alone code.
I used to study physics, and I found that a great deal of innovation and ideas come from so-called physical intuition. There's another path, of course — logical reasoning, the deep search capability of mathematics.
The history of human scientific research shows that these two things can't substitute for each other. For some problems, mathematical thinking is faster; for others, physical thinking is faster.
So to build true AGI, this part is unavoidable. You can train it into the same model, or into different models in layers. But the bigger point is: this capability must be acquired from the real world, or from simulation, or some other way — and then injected into the AI system.
On the application side, people are also somewhat underestimating the importance of embodied AI.
Recently, both in Silicon Valley and in China, there's been massive FOMO around RSI — both the hype and the fear.
My view is: with current language model capabilities, full RSI can't actually be achieved. Because it's missing the most critical link: physical re-manufacturing (self-replication in the physical world), which language models can't touch.
Also, language models can modify their own training methods and much of the training pipeline, but fundamentally we need to scale up — more compute, more chips, more electricity — and language models can't do that either.
That's why in the United States, data center construction has become a major bottleneck. You still need people dealing with reinforced concrete; you still need people tuning semiconductor fab parameters.
But five years from now, I believe people will realize embodied AI can do these things in the not-so-distant future. That's the day true RSI arrives.
Right now, the digital world is moving far faster than the physical world.
I often joke: if Anthropic fired all its employees today and used agents to develop the next-generation model, progress would definitely slow — human taste still matters — but would it fail to produce the next model? I don't think so; it definitely could. But if a semiconductor fab had no people and relied entirely on AI — that definitely wouldn't work.
Also, today, building an app no longer requires a big team, a long timeline, and tons of coordination the way it did five years ago — like spending a year on one app. Today, one person with enough quota can spin up a bunch of agents and build it in a few hours.
Building an airplane today probably still takes 1,000 people, an enormous supply chain, and many factories. Maybe five years from now, with the help of embodied AI and general language models, building an airplane could be as easy as building an app is today.
At that point, the entire production process closes the loop, and the human economy achieves a true decoupling of the silicon-based from the carbon-based.
Maybe one day we'll still need to carve out a Chinese path for AI development — and embodied AI could be a very important starting point, or a very important part of it. Maybe afterward, every company will train its own language model. Who knows.
🤵🏻♂️ Huazhe Xu (Poke Robotics)
I think what's overhyped is, obviously, the pursuit of data volume.
In the market now, someone says they have a million hours of data, someone says a few million, and someone says so-and-so has already bought this much data — and if you don't buy, you're falling behind.
I've also seen some models on the market trained on truly enormous volumes of data that didn't actually turn out beyond expectations — no better than models trained on tens of thousands of hours. Blindly chasing volume is severely overhyped.
What's undervalued, I think, is the progress of general intelligent models — the progress of embodied AI models.
Whenever people talk about embodied AI solving general problems, they say it might take 5 or 10 years. I think within 3 years, embodied AI can start serving in many places — general-purpose service, not post-trained service.
👦🏻 Koji (Crossing)
Everyone has painted a very exciting future. Maybe even the one Qian Wang mentioned — will building airplanes someday become as easy as building apps?
Thanks to our four guests, and thanks to everyone in the audience for your time. See you at next year's Bund Summit!
