What Will It Take for Robots to Move Beyond Demos? We Talked to AgiBot, AI2Robotics, Luming Robotics, Inverse Matrix, and JD.com
You have to reach for the heights and still make it across the river.
You have to reach for the sky, and also cross the river.

👦🏻 Host: Koji
🥷 Editing: Crossing
🧑🎨 Layout: NCon

"If you could only pick one variable, what is the single biggest constraint stopping robots from moving out of demos and into the real world today?"
On September 9, at the JDDiscovery-2026 JD.com Global Tech Explorer Conference, an industry panel titled "Embodied AI Beyond the Demo: How Real-World Data Becomes Industrial Infrastructure" opened with exactly this pointed question. The guests were founders and industry experts whose companies have already made the leap from demo to industrialization.

The panel was hosted by Koji, with five guests: Maoqing Yao, Co-founder, Co-President, and President of the Embodied AI Business at AgiBot; Peng Zhang, Co-founder of AI2Robotics; Chao Yu, Founder and CEO of Lumos Robotics; Boyuan Chen, Founder of Inverse Matrix; and Haoyang Huang, Director of the Image and Multimodal Lab at JD.com's Explore Academy.
Each represents a different position in the embodied AI value chain, and each is an active practitioner in this wave of embodied AI industrialization.
The discussion began with the most practical question of all: What exactly is still missing for robots to move from demos into the real world? The answers gradually converged in the same direction — the real barrier is no longer just single-point model capability, but data, feedback loops, and system-level capability. And the arrival of GPT-6 Astra pushes the question one step further: will the paradigm of embodied models themselves have to change too? Below is the conversation, edited by Crossing:
What's limiting robots from entering the real world?
👦🏻 Koji
If you could only pick one variable, what do you think is the single biggest thing limiting robots from entering the real world today?
🧑🎨 Maoqing Yao, AgiBot
There are indeed many variables. If I had to pick just one, I'd say the core one is "intelligence" itself.
For robots to handle all kinds of work as intelligently and flexibly as humans — and to do it efficiently, with a high success rate — the core is intelligent capability.

Maoqing Yao | Co-founder, Co-President, and President of the Embodied AI Business at AgiBot
🤵🏻♂️ Peng Zhang, AI2Robotics
At this stage, one thing we focus on intensely is "real deployment."
From AI2Robotics's perspective, what we care about is whether the robot can genuinely generate productive value — whether it can actually stand on its own feet.

Peng Zhang | Co-founder of AI2Robotics
🕵🏻♂️ Chao Yu, Lumos Robotics
The core is the ability to close the task loop in real-world scenarios.
Breaking it down, there are several layers: first, closing the loop on the task itself; second, reliability and generalization within that loop, which involves the model, the intelligence, and the entire infrastructure; third, whether the loop enables continuous iteration.
This is no longer a single-point problem — it's about overall system-level capability.

Chao Yu | Founder and CEO of Lumos Robotics
🧑🏻🚀 Boyuan Chen, Inverse Matrix
I lean toward the "feedback loop."
First, the intelligence on the model side itself.
We can now produce excellent demos in controlled settings, but real-world scenarios are full of noise and continuous change. In the real world, so-called edge cases aren't the long tail — they're the norm.
So our demand for general intelligence is that it can complete arbitrary tasks in real settings like homes and factories. That requires the model to be intelligent enough to go deeper into real scenarios.
Second — and more important — is feeding timely data from real scenarios, including corner case data, back to the model to form a complete feedback loop.

Boyuan Chen | Founder of Inverse Matrix
🧑🏻🏫 Haoyang Huang, JD.com Explore Academy
JD.com's Image and Multimodal Lab has long worked on multimodal foundation models, learning video generation and video understanding from large-scale internet corpora.
Applied to embodied AI, what I think is most lacking is "scalable, annotatable, high-quality real-robot data with action labels."
This is what we currently see as the biggest constraint on embodied AI development — and the reason JD.com is championing first-person manipulation data.

Haoyang Huang | Director of the Image and Multimodal Lab, JD.com Explore Academy
Frontline experience on deployment and data collection
👦🏻 Koji
Speaking of deployment, Lumos Robotics has quite a bit of industrial deployment practice. Any new observations you can share recently?
🕵🏻♂️ Chao Yu, Lumos Robotics
In practice, we've seen that real-world deployment really does demand a lot of system-level capability.
That includes data, data feedback — how the returned data gets reused in the model — and how that data is mixed in multi-source ratios with earlier data, and so on.
Also, if you rely entirely on manual effort, the pace is extremely slow. So you actually need something like a "Codex for embodied AI" system.
Right now, Lumos is building its own NexCore system as the vehicle for achieving task-loop capability.
👦🏻 Koji
Could you walk us through a concrete case of how this system works?
🕵🏻♂️ Chao Yu, Lumos Robotics
For example, early on we did flexible quality inspection and screw fastening.
We started by collecting data in our own factory, then did model training, then deployed the robots to real production lines. Through that process we ran multiple data iterations — adjusting the returned data based on actual operations, and supplementing and adjusting based on failure cases.
As the robots accumulate operating time and working hours, overall performance metrics start climbing steadily. That's the process we're going through right now.

👦🏻 Koji
Haoyang, from the perspective of JD.com Explore Academy's Image and Multimodal Lab, how do you view the different data collection paths?
🧑🏻🏫 Haoyang Huang, JD.com Explore Academy
Each has its own characteristics.
The first is real-robot teleoperation data. It's the highest-quality data — it faithfully reflects actions and state changes and can genuinely push the model forward — but the cost is extremely high, making it hard to scale.
The second is UMI data, which establishes the mapping between human motions and robot motions at a relatively low cost.
The third is simulation data. The advantage here is that it can cover many long-tail cases — extreme situations that are hard to trigger in the real world.
The fourth — and the one we value most — is egocentric data: first-person-view manipulation data, which can scale much faster.
In our view, these four types of data form a data pyramid:
First, use egocentric or simulation data for model pre-training, so the model learns to understand space and actions; then use UMI data to bridge the connection between human manipulation and robot manipulation; and finally use SFT post-training for fine-grained real-robot manipulation.

Haoyang Huang | Director of the Image and Multimodal Lab, JD.com Explore Academy
👦🏻 Koji
JD.com originally trained its multimodal models on internet data. Now you're combining that with embodied AI and gaining real-robot data from real scenarios. You're training multimodal models in both cases, but what's the same and what's different between these two stages?
🧑🏻🏫 Haoyang Huang, JD.com Explore Academy
One focuses on observation itself; the other directly supervises the world.
Internet data provides large-scale image, text, and even video data, letting the model quickly understand the world. It's more about "supervision for understanding the world" — focusing on observation itself.
But embodied data is different. It's more about "supervision for acting on the world." For example, I perform an action, and that action changes the state of the world in some way — that's the natural gap between the two.
Take an example: a person picking up a water cup.
For a video generation or understanding model, it finds an average representation from a massive corpus and generates a video of someone picking up a cup. It only cares about whether the video is plausible. But for embodied AI, what we care about is how to pick up the cup, from which position, and with how much force. That requires building a closed loop from action to feedback. That's also why JD.com is championing first-person manipulation data.
This morning, President Cao Peng (Chairman of JD.com's Technology Committee and President of JD Cloud) also mentioned that we plan to release 10 million hours of human manipulation data over the next two years, along with over 1 million hours of real-robot manipulation data.
This data all comes from JD.com's real-world scenarios — retail, logistics, home services, industrial — featuring real objects and real people, and it also covers a large number of failure cases during manipulation.
In fact, the biggest problem with simulation is that sim-to-real transfer is very hard. But through real-world collection, you can cover these scenarios.
So from our model R&D perspective, we no longer just focus on how well the model understands the world. More importantly, it's about "building a closed loop from understanding the world, to execution, to feedback, and to the next round of execution."
That's also our shift from generative models to world models, and then to World-Action Models.
Key points on the road from demo to deployment
👦🏻 Koji
In the first question, Mr. Yao mentioned that the biggest gap for robots going from demo to real deployment is "intelligence." From AgiBot's perspective, what's the most important breakthrough needed on the "intelligence" front?
🧑🎨 Maoqing Yao, AgiBot
From a deployment standpoint, I think three things matter most.
First, long-horizon memory — the ability to make dynamic decisions based on the process.
Second, dexterity and precision. The human body is an exquisitely engineered construct — from degrees of freedom and whole-body coordination to external perception and tactile feedback, it's an extremely precise system. Robots are still in catch-up mode; many things that look simple to us are very hard for robots.
Third, error-correction capability. Machines can now often fit certain states and environments, but when the unexpected happens or a failure occurs, it's very hard for them to close the loop and self-correct.
The industry's original approach was to train an increasingly complete end-to-end model or world model — using demonstration data and RL exploration data — hoping a single model would solve everything.
But that idea is fairly idealistic.
Recently, people have been discussing: should we instead use a layered architecture that leverages multimodal large models to handle upper-level instruction understanding, task decomposition, state monitoring, and replanning and error correction after failure — forming an agent harness system?
👦🏻 Koji
From AI2Robotics's perspective, Mr. Zhang, how do you evaluate how far a capability needs to go at the demo stage before it can enter scaled deployment? What signals do you focus on?
🤵🏻♂️ Peng Zhang, AI2Robotics
Between demo and large-scale commercialization lies a process of exploration. Getting from demo to scale has two prerequisites:
First, a demo showcases technical capability — the goal is to verify that the technical framework and approach work in a given scenario. But to generate real commercial value, you need to account for the customer, the variables of the actual deployment environment, how the product integrates with the customer's systems and business model — and truly create value for the customer.
Second, once you have commercial value, you also need sustained, scalable value creation — meaning demand for large-scale deployment. That involves not just technology and model capability, but also closed-loop capability, whether the hardware body can support long-duration stable operation, and whether after-sales, production, and quality systems can keep up.
Over the past few months, in many industrial and new-retail scenarios, we've already achieved long-duration, stable operation across numerous settings.
That's what we see as the prerequisite for scaled commercial deployment.

A foundation model for the real physical world
👦🏻 Koji
Inverse Matrix is the youngest company — and founder — on stage today. What has Inverse Matrix been working on specifically over the past period? Any new results?
🧑🏻🚀 Boyuan Chen, Inverse Matrix
In one sentence: we're building a foundation model for the real physical world.
We've found that the real physical world is highly structured in industrial settings but unstructured in home settings — with very high generalization requirements.
That reminded us of the development path of language models: in 2023 and 2024 there were still vertical-domain LLMs, but by 2026, one general-purpose language model had one-for-all solved those vertical scenarios. The reason is that a general foundation model can leverage generalizable, transferable knowledge and structure across different domains.
We believe the physical world will follow the same path.
What we want is to build a physical-world foundation model that understands human physical intuition and instinct.
It can all-for-one leverage data from different real physical scenarios — home services, retail, industrial — as well as synthetic simulation data, to help the model better understand underlying physical laws.
Recently we've also discovered some interesting phenomena, like the "less is more" effect unique to foundation models.
The same foundation model, with only a small amount of fine-tuning data, can generalize and transfer across different robot embodiments — dexterous hands, wheeled platforms, and more.
👦🏻 Koji
Building a world model that understands human physical intuition and instinct — what's the hardest part?
🧑🏻🚀 Boyuan Chen, Inverse Matrix
Mainly two things.
First, the paradigm itself. Whether people talk about world models or VLA, the core question is how we learn from real physical-world data.
Take egocentric data: people noticed the 1 million hours released before the Dyna-2 model, and Skild AI directly stitching human demonstrations in as context. But what's more noteworthy is that Skild AI spent 3x the cost of data collection on data quality.
For egocentric data, the effective source is action labels beyond states — which actions cause which state transitions in the world. That's the source of our understanding of physical instinct and intuition. Synthetic data is more about providing something that can scale massively.
Second is how we distill, from real physical-world data, the physical structural paradigm the model should learn.
We want the model, after seeing a cup picked up in daytime and at night, to understand that the underlying principle is the same — because whether you can pick up the cup depends only on force, contact, and material. From there, the model can go on to understand complex scenarios like welding and gluing.
The difficulty here is that the real world is noisy and only partially observable. It's not like language models, which can complete all the information from context.
Real physical geometry, materials, constraints, and forces are all hidden variables. How the model extracts these hidden variables from raw data is a critical step.
Together, these two points lead to a judgment:
The paradigm for physical AI needs to move from pixel space and language space back to a higher-dimensional latent space. And the next challenge is data — in other words, the feedback loop.
In fact, our understanding of physics comes from interaction, not from more perception. From age 1 to 10, a person interacts 10 hours a day, 365 days a year — over 30,000 hours — and learns a great deal of physics. Essentially, beyond perceiving the world, we constantly interact with it, continuously verifying and updating our hypotheses in the process.
For example, after seeing every video where cups sit on tables, I might assume the table has suction. But when I knock a cup to the floor, I discover it doesn't — that's where the understanding of physics and causality comes from.
So we emphasize two things: first, at the paradigm level, return to a level higher than pixel space and language space — the latent space — to compress and distill these abstractable, generalizable physical laws; second, apply and interact in real scenarios, using reinforcement learning to form a complete loop.

Boyuan Chen | Founder of Inverse Matrix
In our exploration, we've also discovered several interesting phenomena:
First, a physical AI foundation model has an all-for-one and one-for-all relationship — a small amount of fine-tuning adapts it to different embodiments and different noise environments.
Second, the diversity of the data itself matters a lot, especially multimodal feedback related to force and touch — including data from our collaborations with JD.com in home services, health, and factories — which helps the foundation model iterate quickly.
Third, beyond scaling data and parameters, we've found a third axis — the interaction paradigm and interaction density itself.
We natively introduce actions into the model's prediction and learning. As the density with which the model learns state changes causing state transitions increases, the model continuously develops a better understanding of the physical world.
This also matches one of our intuitions: the higher the density of trial and error, the deeper our understanding of the physical world.
Lessons from GPT-6 Astra
👦🏻 Koji
A couple of days ago, after GPT-6 was released, many creators online used Astra for agentic task understanding and decomposition. What were your impressions after watching those videos?
🧑🎨 Maoqing Yao, AgiBot
I'm increasingly convinced that today's multimodal large models can already handle long-horizon task understanding, planning, and in-process reasoning quite well at the upper-system level. But for more fine-grained manipulation, you still need to solve it at the policy level.
I think a question worth studying next is:
Should the architecture itself be layered, with communication between layers — or will it eventually move toward a single system that adaptively switches between "fast and slow systems" through something like a MoE architecture?
Just as language models used to require people to manually choose between fast thinking and deep thinking, many products now switch automatically. It's undeniable that the latest frontier VLMs are increasingly capable of this kind of reasoning.

Maoqing Yao | Co-founder, Co-President, and President of the Embodied AI Business at AgiBot
Specifically regarding GPT Astra, here are my takeaways:
First, language is a highly efficient "compression of intelligence."
Looking at recent language models and video models, language as a carrier has compressed information and intelligence over the long course of human evolution. And a native large model — even without much post-training or fine-tuning on the physical world or embodied AI — can already direct a closed-loop system to complete tasks quite well.
Second, video generation models prove that scaling laws also hold in high-dimensional spaces.
This year's mainstream video generation models have made a qualitative leap over last year's. Why such a big jump this year?
Because data volume reached tens of millions of hours, and model parameters went from dense models of a few tens of billions to MoE models of one or two hundred billion. With larger parameters and more data, as long as you get the details right, you can indeed extend scaling laws like those of language.
Third, the real moat is data, not the model.
We've seen several leading companies spend thousands of people on cold-start efforts, labeling hundreds of thousands of hours of data to train caption models. We believe that in the low-dimensional spaces of embodied AI — a dozen or a few dozen dimensions — as long as data labeling is done equally well, there will be corresponding scaling laws to extend.
Recently we've also exchanged views with people working on video models, and everyone feels this deeply — data really is becoming more and more important, because the model itself isn't that strong a moat. But data can buy you a lead of 200 to 300 days.
🧑🏻🚀 Boyuan Chen, Inverse Matrix
Astra represents the highest-level form of true VLA — a language model entering the physical world and forming a real physical loop.
But recently I also tried using GPT-6 to operate a gripper to pick up a cup, and it took maybe 40 minutes to an hour to slowly get it up. That's a real case.
This prompted further reflection: where exactly is the ceiling of language models?
We can now see language models operating real robots — Claude has also released a physical AI version of the MCP protocol that can control hardware bodies. But in reality, language models have many ceilings.
Humans have motor-perceptual intelligence and linguistic-cognitive intelligence. Language is indeed a great carrier, but it also caps how well we can describe tasks. For example, I can imagine how to tie shoelaces, but I can't describe it in language.
Right now it can do some simple grasping, but when you get to the real physical world, with requirements on precision, speed, and generalization, it still fails.
On top of that, the reasoning paradigm of language models locks the application ceiling.
Language models still generate token by token. 100 milliseconds is imperceptible in a chat window, but on a real robot that might be three control cycles. The real physical world changes every moment — you can't have a model that fails to respond in time when a new challenge arrives before the action is even finished.
So Astra's arrival was a big shock to me — it showed us what a language model looks like entering the real physical world.
But the core problems and challenges of language models can't be broken through, and that also gives us firmer confidence:
Physical AI will definitely have its own foundation model — one that needs to learn the paradigms of human manipulation, physical intuition, and instinct, and make decisions and close loops in the real physical world more efficiently and quickly.
🕵🏻♂️ Chao Yu, Lumos Robotics
I don't think GPT-6's Astra is purely the ceiling of a pure language model.
From some of our experiments, if you give it a good representation system plus an execution system, its ceiling is much higher than it is now.
For example, in earlier experiments tuning hardware, feeding trajectories in as raw numbers produced very different results compared to feeding in oscilloscope-like comparison images. So I believe that once you equip it (the model) with the right representation and execution systems — whether at the signal level or the execution level — the overall ceiling will be much higher than today.
Based on this judgment, I think the paradigm of embodied models may change significantly going forward — it might even evolve into a subfield of this class of models.
On top of that, it will substantially disrupt the current data paradigm for embodied AI.
Much of the conventional data being collected today — especially data where contact information isn't fully captured — may see its paradigm change dramatically.
Chao Yu | Founder and CEO of Lumos Robotics
👦🏻 Koji
In the embodied AI field, what new work from the past few months do you think is most worth paying attention to?
🧑🏻🏫 Haoyang Huang, JD.com Explore Academy
I'd like to talk about the JoyAI-RA model JD.com released earlier.
We collected a lot of first-person manipulation data and verified a very interesting conclusion: scaling laws on real human-robot manipulation data do exist. Previously we weren't sure how much data it would take to unlock generalization in embodied models. Now we've found that on tasks the model has seen, the metrics genuinely keep improving.
This is consistent with our experience building multimodal models. From massive data, we've found similar scaling laws. Going forward, with even more data, models will be able to learn many of the ways physical laws vary.
This also indirectly shows we need to acquire more real-robot manipulation data.
Right now we're in a gradual transition: moving from predominantly egocentric and simulation data toward real-scenario data.
🤵🏻♂️ Peng Zhang, AI2Robotics
The whole industry, us included, is watching the development of intelligence itself very closely. We think about it from these principles:
First, from a technology-roadmap perspective, which direction should we go.
Second, from a product and deployment perspective, which factors to consider. Because the end goal is the same — we want the technical architecture and roadmap we adopt, and the path by which the robot keeps evolving through closed loops, to achieve the ultimate goal we're aiming for.
Right now everyone is still exploring; no definitive path has emerged dictating how things must be done.
And that's exactly why this round of the embodied AI revolution is a rare moment when China and the United States start from the same starting line — we can explore it together.
Maoqing Yao | Co-founder, Co-President, and President of the Embodied AI Business at AgiBot (left); Peng Zhang | Co-founder of AI2Robotics (right)
We've talked a lot about Astra, but beyond Astra, what has really struck me these past two months is Tesla's Cybercab.
In the US market, cars designed specifically for robotaxis — no steering wheel, no rearview mirrors — are now actually operating on the roads.
Elon Musk has talked about this goal for many years — hoping that one day your car can go out and take rides on its own during idle time — but he also spent many years preparing for it.
Starting around 2014 or 2015, Tesla committed to a pure-vision approach and released end-to-end FSD. During this period, its product line kept narrowing — sending more data from its cars back to data centers rather than building more kinds of cars. In other words, what Tesla has been doing all these years is harnessing data from every Tesla in the world to effectively evolve FSD, getting infinitely closer to that tipping point.
It's accumulating data inside a complete commercial loop, advancing toward its ideal goal.
That's something worth learning from: how do we combine technological development, commercial closed loops, and scaled deployment? How should AI products, AI systems, and AI commercial loops actually be built?
Closing thoughts
We always overestimate the change that will occur in the next two years and underestimate the change that will occur in the next ten. For practitioners in the embodied AI industry, a few messages from today's conversation are especially worth noting:
First, the center of competition is shifting from "model capability" to "system capability."
Embodied AI deployment isn't a single-point technical breakthrough — it's a massive systems engineering effort spanning hardware bodies, models, data, control, and after-sales. Whoever gets this system running — and running fast — gets a ticket to the next round.
Second, the durability of data moats is being re-evaluated.
The lead time on models is shrinking, but a loop that continuously generates data in real scenarios, continuously iterates the model, and continuously produces commercial value — that's where the real gap will open up.
Third, the ceiling of language foundation models represented by Astra may be far from truly reached.
As Chao Yu of Lumos Robotics put it, the paradigm of embodied models may change dramatically — and may even evolve into a subfield of this class of models.
