Is OpenAI Astra a Crushing Blow to Embodied Models? Seven Takes from Two Roundtables
After Astra's Debut, Do Embodied AI Companies Still Have a Moat?
After Astra's Stunning Debut, Do Embodied AI Companies Still Have a Moat?


Recently, I moderated two forums on Embodied Artificial Intelligence back to back: the Inclusion Conference on the Bund on September 10, and the JDD JD.com Global Technology Explorer Conference on September 9.
The timing coincided with the buzz GPT-6 Astra generated in the embodied AI space. Across social media, people were experimenting with getting it to control robotic arms and grasp objects, and the discussions followed.
So at both roundtables, we ended up talking about the same question:
What kind of shock does GPT-6 Astra deliver to embodied models? How much independent space will embodied models retain? Are general-purpose models the future?
2026 Inclusion Conference on the Bund | September 10, 2026
JDDiscovery-2026 JD.com Global Technology Explorer Conference | September 9, 2026
The panelists offered many brilliant — and quite different — perspectives.
I wanted to bring them together. So I asked Astra to score each of the seven panelists' on-stage remarks on a "General-Dominance Index." It's a 0–10 scale of where they stand:
• 10: General-purpose models are the future. As general capabilities keep expanding, the room for standalone embodied models will shrink dramatically — they may not even need to exist separately.
• 0: Embodied models will remain fully independent. Physical intelligence needs its own foundation model, and that independent space won't be swallowed by general models.
🚥
This index only reflects where someone's position leans. It's not a judgment of who's right or wrong.
It's Astra's interpretation of the on-stage remarks, not the panelists' own self-assessment.
Below, I've re-organized the seven panelists' relevant remarks, ordered from lowest to highest index score.
Boyuan Chen | Founder, Inverse Matrix Tech
Boyuan Chen, Founder of Inverse Matrix Tech | JDDiscovery-2026 JD.com Global Technology Explorer Conference | September 9, 2026
General-Dominance Index: 1/10 (a reference score assigned by Astra based on his remarks — for reference only)
A strong advocate for an independent physical foundation model Astra represents the highest form of a true VLA — a language model entering the physical world and forming a real physical closed loop.
That said, I recently tried using GPT-6 myself to operate a gripper to pick up a cup. It took maybe 40 minutes to an hour to slowly lift it. That's a real case.
It pushed me to reflect further: where exactly is the ceiling of language models?
We can now see language models operating real machines, and Claude has released a physical AI version of the MCP protocol that can control hardware. But in reality, language models have many ceilings.
Humans have kinesthetic-motor intelligence and linguistic-cognitive intelligence. Language is a great carrier, but it also caps how well we can describe tasks. For example, I can imagine how to tie shoelaces, but I can't describe it in words.
Right now it can do some simple grasping, but in the real physical world — where precision, speed, and generalization are required — it still fails.
On top of that, the language model's reasoning paradigm locks in the application ceiling.
Because language models still generate token by token, 100 milliseconds means nothing in a chat window, but on a real machine that could be three control cycles. The real physical world is changing constantly. You can't have a situation where the motion isn't finished executing, a new challenge arrives, and the model can't respond in time.
So Astra's arrival was a real shock to me. It showed us what a language model can do when it enters the physical world.
But the core problems and challenges of language models are also impossible to break through — and that gives us even firmer confidence:
Physical AI will definitely have its own foundation model. It needs to learn the paradigms of human operation, physical intuition, and instinct, so it can make decisions and close the loop in the real physical world more efficiently and faster.
Qian Wang | Founder & CEO, Independent Variable Robotics
Qian Wang, Founder & CEO of Independent Variable Robotics | 2026 Inclusion Conference on the Bund | September 10, 2026
General-Dominance Index: 2/10 (a reference score assigned by Astra based on his remarks — for reference only)
Embodiment needs dedicated models; general capabilities can be borrowed This really is the hottest topic in the industry. Just this morning I saw a video of someone using Astra to control a dexterous hand.
This version of the model was actually trained on a large amount of robot data. Since early last year, OpenAI and Anthropic have been purchasing robot data in bulk, and they also run their own small-scale data factories collecting robot data. So this isn't particularly surprising — it was expected.
I've always had a "provocative take" that colleagues working on the front lines of large models should endorse:
Today, the substance of intelligence lives in the data — and part of it lives in evaluation. The model is more like a distiller and a container: during training it's a distiller, and during deployment it's a container.
When building a frontier AI system today, the model's role has actually become extremely small. The bigger factor is data — you need ever-better data. How you create data, select data, validate data, annotate data — these are exactly the things advancing AI systems.
So the question becomes: if OpenAI or Anthropic wants to do embodied AI, do they still need the same data infrastructure, validation infrastructure, and other infra we rely on? Do they still need to collect data from the real world, build simulators, run real-world evaluation, integrate with hardware, and so on?
If they still have to do all these things, then they're just another embodied AI company. Because none of these difficulties are reduced by their advantage in foundation models. If anything, embodied AI companies may have a substantive advantage.
What everyone is really worried about is: can a foundation model generalize from general intelligence in other domains into embodied intelligence and deliver a "dimensional-strike" blow?
I don't think that's likely.
First, this has never happened in the history of language models. Recently it's been the same with 3D, and with computer use — even though today's video generation models are excellent, that hasn't directly translated into a natural advantage at the action level or the robot-control level.
But the flip side is instructive: pretraining really does work.
Many people in the industry have worried that all the pretraining we do, with massive amounts of general data mixed in, might actually be harmful — might hurt model performance. With Astra, we see it's at least a bonus, at least not a bad thing. We should keep going down this path with conviction.
Finally, the only gap between us and Anthropic or OpenAI might be this: they fold all this infra into a general-purpose or language model, while we're only building specialized models for a specific industry.
In fact, general-purpose and language models simply cannot run inference and deployment on robots at reasonable speed. We genuinely need a dedicated embodied model to do this — it's the more efficient approach.
Yujun Shen | Chief Scientist, Ant Lingbo Technology
Yujun Shen, Chief Scientist of Ant Lingbo Technology | 2026 Inclusion Conference on the Bund | September 10, 2026
General-Dominance Index: 4/10 (a reference score assigned by Astra based on his remarks — for reference only)
Coexistence through collaboration, with more weight on an independent embodied paradigm My view is quite close to Huazhe Xu's, maybe slightly more radical. I even feel that embodiment will gradually influence the development path of large models — or cloud models.
Going back to how an embodied model ultimately works: it has to be deployed on the robot. What the robot does depends on all the sensors on its body, and those sensors are feeding in a continuous stream of input, even while the robot is operating.
That's different from the turn-based conversational model of the digital world. In the digital world, you ask it a question and it answers — while it's answering, it's hard for you to keep asking. It's linear. But a robot's working environment is a continuous, uninterruptible stream of input. Even while it's reasoning, input keeps coming in, and it has to revise its reasoning at any moment based on new input.
In the end, the embodied model may become a tool used by higher-order models.
As a tool, it may need its own training paradigm, because it has to face that scenario: continuous input streaming in, policy streaming out. That's one of my points.
Why do I say embodied scenarios might open up more development directions for higher-level models — interactive models like Astra?
Here's an example. Conversation between people is actually multimodal — not just language and vision, but many external factors. Say I'm chatting with someone and it suddenly starts raining outside. Seeing the rain might change my mood and the way I talk. Or I notice the other person looks upset, so I drop much of what I was about to say and immediately switch to something else.
Many cloud-based digital-world models haven't accounted for this yet — they don't need to right now. They live behind a screen, and people only interact with them through vision and language.
But if robots are really going to converse with people, the amount of information they can take in should be enormous, and much of it will come from real-world sensors. We observe the weather, sense the temperature — all of this affects the model's output.
If robots really do gradually enter daily life and bring in more data from the physical world and real life, that data could become new corpus for more advanced large language models and model companies, making them more intelligent.
The model would no longer be a cold agent hiding behind a screen. Instead, it could sense its surroundings like a person and adjust how it talks to you as things change.
And that's exactly why, when it comes to training paradigms, I think we may be overestimating "the technology of training models."
I've heard voices in the industry saying embodied AI is about to get dimensionally crushed by today's digital-world models — as if you could just stockpile a batch of physical-world data, apply the previous training methods, and train an embodied model. I don't buy that.
I believe embodiment may follow the same pattern as the digital world, but it definitely won't take the same technical route.
Because the final application scenario determines what the model must be. The digital world is conversational, but embodiment is necessarily a robot reasoning while it works.
Embodied AI will definitely have its own new training paradigm. One that supports continuous processing of physical-world information — possibly 7×24, even 365×24, processing data nonstop, all year round.
Zheng Han | Co-founder & CEO, Sudu Tech
Zheng Han, Co-founder & CEO of Sudu Tech | 2026 Inclusion Conference on the Bund | September 10, 2026
General-Dominance Index: 5/10 (a reference score assigned by Astra based on his remarks — for reference only)
Layered architecture, with general and embodied models carrying equal weight This is actually getting closer and closer to a point I've been making for a long time: the upper and lower layers will separate. The abstract understanding at the top will form its own layer.
I've been saying this for years: Facebook AI Research (FAIR)'s biggest rival is actually OpenAI — and recently Anthropic as well. Another Robotics researcher from DeepMind just joined Anthropic too.
One more thing: the foundation of robotics is skills. The generalization and high success rate of low-level skills must match the upper-level model. You can't have low success rates with some generalization thrown in — that's a fatal problem at the junction between layers.
So when embodied companies interface with upper-level models, they must pay attention to this: generalization at high success rates.
Late last year, we did a demo with the Qwen model handling language interaction. Through its understanding of the surrounding environment and objects, including the abstract instructions given to it, it broke everything down into reasoning-generated operating steps — but it relied on extremely reliable and stable low-level manipulation of the environment and objects to complete the whole sequence of actions.
This year the idea has come back to the forefront, and I think it's the general trend — especially after Astra's arrival.
Huazhe Xu | Founder, Poke Robotics; Assistant Professor, Institute for Interdisciplinary Information Sciences, Tsinghua University
Huazhe Xu, Founder of Poke Robotics; Assistant Professor, Institute for Interdisciplinary Information Sciences, Tsinghua University | 2026 Inclusion Conference on the Bund | September 10, 2026
General-Dominance Index: 6/10 (a reference score assigned by Astra based on his remarks — for reference only)
General capabilities take the lead; physical models remain critical Astra is genuinely stunning. Internally, we tested it the very day it became available.
If you break embodied tasks down into semantic generalization, spatial generalization, and physical generalization, Astra is relatively weak on physics, and relatively strong on semantics and space.
We gave Astra a prompt and it recognized everything on the desk — no matter the category — and it would definitely head to the right place. With previous VLM or large models, the robot would basically get there and start flailing around, and that was the end of it.
But one thing that makes Astra distinctive is that it can actually grasp — and grasp very hard-to-grasp objects. For example, a chopstick lying flat on a table: for a robot, such a small object usually has to be pushed against the table edge, even sliding, before it can be picked up. Astra can pick it up now.
After testing that far, we wondered: would rotation be a weakness? So we tried getting it to stand a chopstick upright and insert it into a cup. It could do that too. So not just 2D space — across all of 3D space, it performs well. But when we asked it to do complex tasks involving physical contact — like using chopsticks to pry something apart — it started to struggle.
I think this is also the last stronghold of embodied robotics companies: understanding physical change.
We also tried getting Astra to stack boxes and fold clothes — things embodied companies love to do — and Astra couldn't do any of it. And among all extremely complex tasks, making mapo tofu would surely be one of the most complex.
Back to Astra: what it can do now is still embodied tasks that are relatively spatial-semantic in nature. Among the most complex contact-rich tasks, the ones it can manage are putting a chopstick into a cup, and handover — picking something up, passing it to the other hand, and placing it on the other side. We tested many times and it succeeded once or twice.
What Astra suggests to me is that in the future these things should be strung together — a bit like what Professor Mengdi Wang (Director of the AI Innovation Center at Princeton University) said in her keynote just now: embodied models should be hooked into Codex, or into a harness environment.
On one hand, embodied models need to work generally; on the other, the harness needs strong semantic capability. Take grasping fruit: fruit from the market rots, so embodied models usually start with toy fruit or foam fruit — the model can't tell the difference. Something like Astra can provide feedback, laying out some CoT about what happens after you grasp, and then act.
I believe the end state should be a complete system — large model, embodied model, and physical world connected together — to truly solve the problem.
Maoqing Yao | Co-founder, Co-President, and President of the Embodied Business Group, AgiBot
Maoqing Yao, Co-founder, Co-President, and President of the Embodied Business Group at AgiBot | JDDiscovery-2026 JD.com Global Technology Explorer Conference | September 9, 2026
General-Dominance Index: 7/10 (a reference score assigned by Astra based on his remarks — for reference only)
Leaning toward a unified model, while keeping a dedicated execution layer I'm increasingly convinced that today's multimodal large models, at the upper system level, can already handle long-horizon task understanding, planning, and in-process thinking much better. But for finer-grained manipulation, you genuinely still need the policy layer to solve it.
I think a question worth researching next is:
Is the architecture itself layered, with communication between layers — or will it evolve toward a single system that adaptively switches between "fast and slow systems" through something like an MoE architecture?
It's like how language models used to require people to manually choose between fast thinking and deep thinking, but many products now switch automatically. There's no denying that the latest frontier VLMs increasingly have this kind of thinking capability.
Specifically, GPT Astra gave me these takeaways:
First, language is a highly efficient "compression of intelligence."
Looking at recent language models and video models, language — as a carrier refined over the long course of human evolution — compresses information and intelligence. Yet a native large model, even without much post-training or fine-tuning on the physical world and embodied intelligence, can already direct a closed-loop system to complete tasks quite well.
Second, video generation models prove that the Scaling Law also holds in high-dimensional spaces.
This year's mainstream video generation models made a qualitative leap over last year's. Why such a big jump this year?
Because data volume reached tens of millions of hours, and model parameters went from dense models of a few tens of billions to MoE models of one or two hundred billion. With larger parameters and more data, as long as you get the details right, you really can extend a Scaling Law like language's.
Third, the real moat is in data, not the model.
We've seen several leading companies put thousands of people into a cold start, annotating hundreds of thousands of hours of data to train caption models. We believe that in the low-dimensional space of embodied intelligence — a dozen or a few dozen dimensions — as long as data annotation is done equally well, there must be a corresponding Scaling Law to extend.
Recently we also talked with people working on video models, and the shared realization runs deep — data really is becoming more and more important, because the model itself doesn't have that strong a moat. But data can buy you a 200-to-300-day lead.
Chao Yu | Founder & CEO, Lumine Robotics
Chao Yu, Founder & CEO of Lumine Robotics | JDDiscovery-2026 JD.com Global Technology Explorer Conference | September 9, 2026
General-Dominance Index: 8/10 (a reference score assigned by Astra based on his remarks — for reference only)
General models dominate; embodiment may become a subfield I don't think GPT-6's Astra is simply the ceiling of a pure language model.
From some practical experience, if you supplement it with a good representation system plus an execution system, its ceiling is much higher than where it is now.
For example, in earlier experiments when we were tuning hardware, feeding in trajectories as pure numerical values produced very different results from feeding in oscilloscope-style comparison images.
So I believe that once it's equipped with suitable representation and execution systems — at either the signal level or the execution level — its overall ceiling will be much higher than it is now.
Based on this judgment, I think the paradigm for embodied models could change dramatically going forward — it might even evolve into a subfield of this kind of model.
Also, it will significantly disrupt the current data paradigm for embodiment.
Much of the conventional data being collected now — especially data where contact information is insufficient — that whole data paradigm may change substantially.
Afterword
As you can see, the discussion Astra has sparked is rich and diverse.
Standing at this crossroads for embodied intelligence, I don't think the future is necessarily a single-choice question between "general" and "specialized." What really matters is who can connect semantics, space, physics, and execution into one complete system.
I hope discussions like this keep multiplying, and Crossing will continue to initiate and take part in these important conversations.
Recommended Reading




Join the Member Community Crossing has always been committed to being the connective hub for founders in this golden age of AI — a platform where you meet, exchange ideas, and make decisions at critical moments. Founders, PMs, developers, and super individuals who take action are welcome to join our WeChat member community.

After joining, you'll get:
- Deep participation in AI R&D and product experience exchanges
- Direct access to leading large-model and cloud service providers
- Media exposure opportunities on the Crossing platform
- Fast-tracked connections to top AI investors
- Priority access to AI Hacker House (Shanghai Caohejing) events and services