Independent Variable Robotics' Qian Wang: China will produce the world's top embodied AI company, aiming to be the OpenAI of embodied intelligence | Founders Talk
Lands the Year's Biggest Robot Funding Round to Date
Unity Ventures portfolio company Independent Variable Robotics recently announced the completion of a 1 billion yuan A++ funding round, jointly invested by ByteDance, HSG, Beijing Information Industry Development Fund, and Shenzhen Capital Group, making it the only embodied AI company in China backed simultaneously by all three major tech giants — Meituan, Alibaba, and ByteDance.
This largest robotics financing deal of the year has thrust founder Qian Wang into the spotlight. He is among the most fervent advocates in the industry for end-to-end embodied physical models. Before this technical approach became market consensus, Wang had already convinced early investors like Unity Ventures with his unique understanding of technology and his unwavering dedication to AI.
In a recent conversation with LatePost, Wang shared the journey of an atypical embodied AI entrepreneur and where he is headed. Wang wants to do what OpenAI does — original innovation from zero to one. Opportunities to change the world belong to entrepreneurs like him.
This article is republished with authorization from LatePost (ID: postlate); author: Shen Yuan, editor: Song Wei

The first interview question took Wang 30 minutes to answer. It took so long because his experience is extraordinarily complex: undergraduate studies in Tsinghua University's Department of Electronic Engineering, graduate studies in biomedical engineering, a PhD in Robotics Learning at USC, and a first job running his own quantitative fund.
In summary, he is an atypical embodied AI entrepreneur: he has neither experience working at major Chinese or American tech companies, nor prestigious academic titles.
This does not diminish Wang's confidence.
During the interview, Wang rarely hesitated. He generally spoke rapidly and got straight to the point, while citing extensively to explain why others couldn't succeed, and why he could.
Wang was already working on neural networks in 2009. The architecture he designed was one step away from the transformer — what he describes in his own words as a Turing Award-level miss, and also the starting point of his technical confidence. He is the person in the embodied AI industry who most enthusiastically embraces end-to-end embodied physical models.
In the early days, this confidence made some investors hesitate, but more and more were convinced. Xinyu Wang, partner at Meituan Longzhu, described Qian Wang as someone with his own unique understanding of technology and persistent judgment. After tracking Wang for a year, Meituan became an important shareholder of Independent Variable Robotics.
On January 12, Independent Variable Robotics announced the completion of its 1 billion yuan A++ funding round, just four months after its previous round. The lead investor this time was ByteDance.
This is someone waiting for an opportunity to change the world. Wang wants to do what OpenAI does — original innovation from zero to one. He wants to be number one.
Missing a Turing Award-Level Piece of Work
LatePost: Before founding Independent Variable Robotics, your previous experience was running a quantitative fund in the United States. How did such a huge transition happen?
Wang: The transition wasn't actually that large, because the underlying technology was the same. My PhD was in Robotics Learning, which was mainly deep learning — quite similar to the tools used in quantitative finance. For someone doing AI who wants to make money, quantitative finance is very direct.
LatePost: How did you first get the idea to do AI?
Wang: When I was young, I mainly wanted to do mathematics and physics. Later I realized that compared to 100 years ago, the professional lifespan of theoretical physicists and mathematicians had become very short. So I wanted to build an engine for human brain intelligence — that is, AI.
LatePost: You were in Tsinghua's electronics department for undergrad, but switched to biomedical engineering for graduate school. Why?
Wang: Why do people fundamentally believe AI can be built? Because there's a natural intelligent system right in front of us — the human brain. But at that time, the technical route for AI was statistical learning, with success rates improving 0.1% per year, and you couldn't even tell if it was due to overfitting. So I thought of neural networks.
At that time, absolutely no one thought neural networks were a good thing. I searched through every lab in Tsinghua's entire School of Information, and not a single professor was working on neural networks. So I went to the biomedical engineering department, mainly researching computational neuroscience.
My advisor had just returned from the United States and told me that Geoffrey Hinton had done something called deep learning. I looked at it and thought, isn't this just neural networks? So I actually started doing deep learning in 2009, among the earliest in China.
LatePost: Many materials say you were among the earliest in China to work on attention mechanisms. How did you find your way to this direction?
Wang: The highest form of human intelligence is self-consciousness, below that is consciousness, and below that? Most people would say attention. So I wondered if I could put it into neural networks and try it out. By 2014, I had the paper done.
Note: The paper is titled "Attentional Neural Network: Feature Selection Using Cognitive Feedback," https://arxiv.org/abs/1411.5140
The paper proposed a new neural network framework that put top-down attention mechanisms and bottom-up feature extraction into a unified model.
This paper was submitted to NIPS (now NeurIPS) and was among the earliest three papers on attention mechanisms. You could say I missed a Turing Award-level piece of work.
At NIPS, there were three papers. The other two were from teams at DeepMind and ETH Zurich. Our architecture was far closer to today's Transformer than theirs.
LatePost: What was the main difference?
Wang: Multiplication operators are inherently very difficult to converge, especially when you stack many layers.
At the time, I was interning at Microsoft Research Asia and had exchanges with Kaiming He, Jian Sun, and others. They were working on ResNet. When Transformer came out, I realized that what we were missing was simply connecting our architecture with ResNet — ResNet makes convergence very easy to stabilize.
LatePost: Why did you switch to robotics after the paper was published?
Wang: I wanted to pursue further studies abroad after my master's. At that time, the first wave of China's "AI Four Dragons" had emerged, but I wasn't very interested in doing security markets. I wanted to find a major direction where AI could truly be applied, so naturally I thought of robotics.
LatePost: Were people in robotics also using deep learning methods at that time?
Wang: In the United States, there were only a few groups doing deep learning in robotics. One was Sergey Levine (co-founder of robotics company Physical Intelligence) and his advisor Pieter Abbeel, whom everyone knows today. There were also groups at MIT and CMU. In the end, I chose USC. So I was classically trained in embodied AI — at the time we still called it Robotics Learning.
LatePost: This direction later hit a wall?
Wang: By 2018 and 2019, the entire AI field felt somewhat stagnant. In robotics, this manifested as deep reinforcement learning hitting a wall, because it inherently had a terrible characteristic: the demand for data grew exponentially with task difficulty. At that time, no one was doing imitation learning either, so the entire direction seemed wrong.
LatePost: What about simulation?
Wang: That doesn't work either. The gap between the physical world and the virtual world is simply too large. The physical world is usually difficult to observe and has extreme randomness.
(Wang pressed his finger against the interview table and pushed it forward.) On one hand, fingers can deform. On the other hand, there's nonlinear friction. When these two things couple together, randomness emerges. This is something you almost cannot simulate. Anything trained in simulation environments cannot be used in the real world. So my final judgment of the entire field was that without some fundamental changes, it might take another thirty or fifty years before robotics could actually work.
LatePost: So you chose to leave academia and start a quantitative fund.
Wang: I was indeed quite depressive at the time. I also didn't particularly like the lifestyle of academia. So naturally I thought I should go make some money, and the most direct way was quantitative finance.
There was also precedent for this. The most typical example is James Simons of Renaissance Technologies. He won the Fields Prize with Shiing-Shen Chern, was very successful in quantitative finance, and then in turn donated money to his alma mater, Stony Brook University, building up its mathematics department very well. There's also an example in the AI field — Wenfeng Liang.
Silver Bullet: GPT-3
LatePost: When did you start thinking about returning to AI and embodied intelligence to start a company?
Wang: In 2021, GPT-3 came out, and I immediately felt this was a huge paradigm shift. Because it had few-shot learning.
People had been pursuing this for decades, and no one had really found it. The biggest problem with reinforcement learning is exponential explosion, but GPT-3 required less and less data to learn a new task. With ChatGPT, zero-shot learning even emerged.
By the way, I find it quite absurd that some people today are reviving reinforcement learning in robotics and calling it a new direction.
LatePost: Why did no one think of the GPT-3 route before?
Wang: This is too counterintuitive to people's instincts. People used to assume by default that specialized models must be the best, but now no specialized model can outperform general models.
This is the Silver Bullet. I originally thought I would have to wait 30 or 50 years, but now there was hope for solving the problem.
LatePost: When you saw GPT-3, did you think back to when you were doing neural networks at Microsoft Research Asia?
Wang: That's exactly why I had to come back and do this. Why from PhD in robotics to quantitative finance, and then back again — this is actually consistent throughout. I simply want to do AI, nothing more. I just changed methods a few times in between.

Image source: WALL·E, the origin of Independent Variable Robotics' model name WALL-A.
China Does Hardware, United States Does Software — Impossible
LatePost: After deciding to build robots, why did you choose China instead of staying in the United States?
Qian Wang: I had considered the US. But after looking around in 2022, I felt the entire hardware ecosystem there had basically collapsed.
Supply chain is the classic problem. If a robotic arm breaks in an American lab, repairs can take two months. In China, it's one day. That's an order-of-magnitude difference.
More importantly, Silicon Valley VCs weren't investing in hardware anymore. Figure AI's early investors were either the founders themselves, or NVIDIA, OpenAI, Microsoft, and Jeff Bezos (Amazon's founder). There were no real financial institutions involved.
The same was true for talent. Silicon Valley isn't short of good hardware engineers, but everyone was at Apple and Meta. No one wanted to leave, or if they did, the goal was to get bought back by Apple.
From talent flow to information flow, capital flow to supply chain flow — the Silicon Valley hardware ecosystem had completely fallen apart.
LatePost: China's advantages are obvious, but what about the disadvantages? Funding, for instance, and compute?
Qian Wang: Fundraising in China is definitely much harder than in the US. But for Embodied Artificial Intelligence, what's limiting scaling up isn't compute — it's data. And data costs in China are clearly an order of magnitude lower than in the US.
So when you do the math: funding is one order of magnitude lower, but costs are also one order of magnitude lower. It basically evens out. And the funding disadvantage isn't permanent, while the cost advantage is sustained.
LatePost: What about human resources?
Qian Wang: In 2022, people still debated Silicon Valley's talent advantage. Now no one asks that anymore, because everyone knows it's the same people working on AI in Silicon Valley and in China. We were all classmates in college.
LatePost: Has your assessment changed since you started the company?
Qian Wang: The US has moved faster than I thought.
Take Figure — one reason its valuation is so high is that it's carrying the whole narrative of manufacturing reshoring to the US. It's really pouring in an order of magnitude more money to aggressively build hardware in-house. Next, it plans to self-produce joints, motors, batteries, even motor winding equipment. The only thing missing is screwing the screws itself.
People used to say China does hardware, the US does software, and in some sense they can stay out of each other's way. That's completely impossible. American companies including Figure are doing hardware no worse than domestic companies. Whether they can mass-produce is another matter, but pre-mass-production hardware quality — I think they're doing better than 99% of domestic companies.
LatePost: When you returned to China to build your team, who did you reach out to first?
Qian Wang: Our CTO, Allan Wang. We met in 2021. His boss at IDEA Research Institute was a co-author on my Attention paper. When I started in quantitative finance, I had a lot of infrastructure work that I hadn't done much of before, and they recommended Allan to me. He got into large models quite early. In 2021, China's two major open-source large model organizations were BAAI and IDEA.
By the way, I believe many embodied intelligence companies today will struggle with the coupling of infrastructure and algorithms, because they haven't done it before. There's still quite a large gap between the two.
When I found Allan, he was painfully working on AI落地 projects, because that's just hard to落地. Even now, if you're not doing coding, you can't落地. After I talked to him, he felt robots were actually the perfect thing to落地. Of course, looking back today,落地 still has many challenges.
LatePost: Not so easy to落地, right?
Qian Wang: Because beyond the model, robots have many other elements — hardware, systems, and so on. But after I told him then, he came to Beijing to find me first, and once he came, he never went back.
The Reason You Can't See Embodied Intelligence's Scaling Law Is Because the Data Is Terrible
LatePost: Independent Variable Robotics' WALL-A model has been described as an end-to-end embodied foundation model, on par with large language models. With such major route divisions in embodied intelligence, why are you so certain about end-to-end?
Qian Wang: When the company was founded at the end of 2023, no one believed in end-to-end. Investors all told me I should still build a hierarchical model or specialized model. But if there was no paradigm shift, if I was still doing specialized or hierarchical models, why would it be my turn to do this? Proprietary models absolutely cannot succeed. You must do foundation models, and then do proprietary models from there.
LatePost: What's the weakness of hierarchical models?
Qian Wang: Say you want to grasp an object. Following the hierarchical approach, you first reconstruct the object's 3D shape, then estimate its center of mass, select a grasp point, generate a trajectory to contact that grasp point, and finally successfully lift the object.
First, 3D reconstruction can't perfectly reproduce an object's surface physical properties — those burrs, pits. It's extremely sensitive to physical contact. So an initial tiny error gets cascaded and amplified very quickly in a hierarchical model. The more layers, the faster the error amplifies.
People followed this route for 80 years and got nothing done.

LatePost: Can end-to-end avoid such problems?
Qian Wang: Because you can backpropagate from the final grasping result to correct the initial grasping motion, trying to increase success rates for certain grasp positions. End-to-end doesn't need 100% perfect reconstruction.
Also, the end-to-end approach didn't originate in the large model era. In 2014-2015, Sergey Levine and others, including us at the time, used end-to-end methods. Around 2018, machines first achieved true general grasping, also using end-to-end deep reinforcement learning.
LatePost: What's the main bottleneck affecting model performance improvement now?
Qian Wang: Data quality is the most important. Some people say they can't see embodied intelligence's scaling law — that's because the data is terrible, all noise.
Previously, 80% of the work was on model algorithms. Now 80% is on data, and for the rest, you want the model to decide what to do as much as possible. This is a major methodological shift.
LatePost: Simulation data doesn't work?
Qian Wang: You need high-quality real data, doing actual tasks in real physical environments.
LatePost: What about virtual simulation environments like NVIDIA's Omniverse?
Qian Wang: GR00T's first version was terrible because it used pure virtual simulation data. Later versions started shifting toward blended data.
I often tell investors this logic: do you believe any simulation company can surpass NVIDIA in compute? NVIDIA caps the ceiling for all these companies, and NVIDIA itself has shifted to real data.
Our generation of PhDs — everyone started with simulation. Now not a single person is still doing simulation, because it just doesn't work.
LatePost: But many people in the embodied field are still doing simulation data.
Qian Wang: I'm a true, orthodox, classically trained roboticist. Some others come from CV or graphics. Maybe they think it's feasible. But we really stepped in all the pits back then.
LatePost: Compute isn't the core bottleneck?
Qian Wang: At least not currently. Under equivalent capability conditions, multimodal models are one to two orders of magnitude smaller than language models. Language models need to memorize a lot of things. Physical world models don't need to memorize much — they just need to know physical laws.
This was also a consideration in my choice to return to China. The embodied field doesn't currently have a compute bottleneck problem.
LatePost: Theoretically, embodied foundation models, like multimodal models, are very difficult to converge.
Qian Wang: Multimodal models are hard to train because data is naturally missing. First, there's a lack of temporally continuous, causal understanding. When a person first sees a cat, they can walk around it — so your understanding has temporal continuity. You also know your own position, so you have a 3D understanding of the cat. Finally, you can interact with it — shake its paw, play for a while. These are all additional information, so humans don't need to see ten thousand pictures of cats to know what a cat is.
When you factor in action continuity, you'll find embodied intelligence models are actually easier than pure multimodal models. Ten years from now, we'll find that the best multimodal models are embodied models. I've told many multimodal researchers: if you really want to do multimodal well, you should come do embodied intelligence.
LatePost: Critics of end-to-end would say — how can walking and solving a Rubik's cube with your hands, two completely different things, be expected from one model?
Qian Wang: First, this really doesn't need to be one model. End-to-end refers to the structure within the model, not functional partitioning. The human brain is also end-to-end, but different regions handle different functions.
However, in practice, putting navigation and manipulation together really does perform better.
LatePost: The model shows more generalization?
Qian Wang: Everything improves a bit. The most typical is COT (chain-of-thought). What people call embodied COT is still doing a language COT first, then attaching a control model — that's still hierarchical.
We were among the earliest worldwide to do native COT, starting at the end of 2024, and around the same time as Gemini Robotics in 2025. Ideally, it can do infinitely long strategies and planning.
LatePost: Can you give an example?
Qian Wang: Say you give it a blueprint with building blocks beside it. It can build according to the blueprint. First, it understands the blueprint. Second, it evaluates the gap between each step and the final result. Third, it actually builds it with its hands.
LatePost: Which part still isn't good enough?
Qian Wang: Overall, everything still isn't good enough. The core reason is insufficient data volume. Algorithms matter too, but data is first.
LatePost: What do you think of Fei-Fei Li's world model?
Qian Wang: Fei-Fei Li's spatial intelligence leans toward 3D generation. But as I said, knowing all 3D shapes doesn't mean you can do everything.
A perfect spatial intelligence model only equals 40% to 50% of a complete embodied intelligence system. The rest is all related to direct physical contact processes.
Hardware Must Be Defined by AI
LatePost: Independent Variable Robotics has released two generations of wheeled robots. Rumor has it you only started building hardware at the end of 2024. Why so late?
Qian Wang: We've always felt AI is first-principles; hardware is second-principles. Early on, hardware conditions weren't very mature, and we were always a small team. Later we realized that after building our own hardware, many AI problems actually became easier to solve. We started large-scale hardware hiring in January 2025.
LatePost: You're from an embodied intelligence background. Didn't you think hardware was important at first?
Qian Wang: A company's resources are limited, especially early on when there wasn't much money. We relied more on suppliers then.
LatePost: You said AI problems became easier after building your own hardware. Can you give an example?
Qian Wang: For example, although they're all robotic arms, whether they're defined AI-natively makes a huge difference. Because I know how robotic arms should be used in data collection and inference stages, and only with this kind of robotic arm naturally suited for AI can you do meaningful research.
Now there are two views. One thinks you should first build perfect hardware, then do AI based on that hardware. This is completely wrong. The other is my view: you must use AI to define hardware.
Another example is the dexterous hand. The human palm has no muscles, so it wraps around objects very well. But many dexterous hands put motors inside, making them thick and stiff, while keeping the external shape human-like. When that happens, the palm loses its function entirely — it can't wrap around anything. When grasping objects, you're actually applying force from the base of the fingers.
This is a classic example of hardware design that only happens at companies that haven't collected data or trained models.
LatePost: Does the capability of your dexterous hand also depend on iterating your embodied physical model?
Qian Wang: The physical laws, motion patterns, and understanding of object properties that the foundation model learns don't change whether you're operating with a gripper or a dexterous hand. If you have a good gripper-based model, training a dexterous hand on top of it saves enormous resources and time.
Of course you still need fine-tuning and post-training, but the principle is similar to large language models — the better you train on English, the easier it is to transfer to Chinese.
LatePost: Elon Musk said dexterous hand technology is harder than Tesla building cars, second only to SpaceX's reusable rockets.
Qian Wang: Hardware is genuinely difficult, but I see hardware and model capability as two parallel tracks. We're also building dexterous hands, but mainly to help train our models.
Actually, most scenarios don't require a hand with the same degrees of freedom as a human. One reason is cost, another is it's not that useful. Humans can perform very complex tasks with just grippers, and grippers are sufficient for at least half of most scenarios.
LatePost: But people feel that achieving a human-like dexterous hand would be a huge breakthrough.
Qian Wang: I'm not so sure. High-DOF dexterous hands are indeed very useful for some tasks, but most of the time they probably just provide emotional value. It looks like a hand, complex and impressive, and that's about it.
LatePost: How far along is your dexterous hand?
Qian Wang: We've built a 20-DOF hand with decent results, but this definitely isn't our main focus — it's more for our model training.
LatePost: Your robots are wheeled rather than bipedal. How did you think about that?
Qian Wang: Legs have two fundamental problems. One is safety — they're inherently more prone to falling than wheeled designs. The other is cost, because the number of motors and joints is an order of magnitude higher than wheeled designs.
LatePost: But don't they have any advantages?
Qian Wang: Their utility isn't that great. Sure, there's emotional value, but setting that aside, how many indoor scenarios actually require legs? The utility doesn't outweigh the disadvantages.
LatePost: So you won't do bipedal robots?
Qian Wang: We might, but we want to do it where it makes sense. When building a company, what you choose not to do is often important, and this is where we've chosen not to.
"We Want to Build a Company Like OpenAI"
LatePost: Some investors say your technical approach hasn't changed from day one, and you can sit tight without rushing to commercialize. That must have made early fundraising difficult.
Qian Wang: Some investors' logic at the time was: you're neither ByteDance nor Google, what makes you think you can build a large model? Even if embodied intelligence needs a large model, why you and not someone else? Many companies had already raised over a billion yuan, and we were still doing our angel round.
LatePost: How did you respond?
Qian Wang: There was really no way to respond. This is a problem with China's capital market — people don't believe technology is first-principles. Subconsciously they think anyone can do technology, that there's no uniqueness.
Because all the past successes were fast followers. There's never been a case of being #1 from the 0-to-1 stage.
LatePost: You believe China can actually achieve #1 in embodied intelligence from the 0-to-1 stage?
Qian Wang: Someone asked me if I want to be the DeepSeek of embodied intelligence. I said DeepSeek is of course a great company, but we want to build a company like OpenAI.
LatePost: Only investors who agree with this point would invest in you?
Qian Wang: Basically everyone who invests in us buys into the logic that we want to be world #1. If you buy into the quick-profit logic, you wouldn't invest in us at all. Some of our shareholders have told me: focus on building the foundation model well, and come to us if you need money.
The two best domestic companies at large models, Alibaba and ByteDance, have both invested in us. We're also ByteDance's only investment in an embodied intelligence company.
LatePost: I heard that in 2024, some investors gave your robot a pop quiz — rolling up toilet paper — and you performed well.
Qian Wang: They gave us three days. They said, don't you claim few-shot learning capability? Then here's a task you've never seen before — do it in three days.
The task was organizing toilet paper. You had to tear off dirty, crumpled parts, apply a plastic sealing sticker, and put it back. It's basically a hotel restroom cleaning workflow.
LatePost: And you succeeded.
Qian Wang: Pretty well. We spent one day collecting data, one day training, and on the third day the investors showed up with piles of various toilet papers. So the actual preparation time was two days.

LatePost: With improved model capabilities, fundraising must be smoother now than in the early days.
Qian Wang: It's better now. For one, people realize China's talent pool and density are in no way inferior to the United States. For another, whether it's DeepSeek or Unitree, people have seen that China can do world-class things, that there's no insurmountable problem. Whether it's resources, compute, or anything else — none of these are fundamental issues.
LatePost: So no one asks why you instead of Google or AgiBot anymore.
Qian Wang: Not really anymore.
LatePost: You seem to have never had those rigid stereotypes from the start.
Qian Wang: Maybe because I roughly understand both sides, China and the US, so I never felt there was anything the US could definitely do that China couldn't.
Team Score: 8 out of 10
LatePost: You didn't have experience managing large teams before. How do you prioritize your time?
Qian Wang: I spend considerable time on recruiting and fundraising. On the technical side, I participate in major technical decisions, and I might personally oversee the most important products.
Most of the time I don't micromanage. If a CEO needs to manage things at that level of detail, something's probably wrong with the company. I'm not a controlling person, and I don't want them coming to me for everything.
LatePost: Compared to other robotics companies, you don't have much halo effect. Is recruiting difficult for you?
Qian Wang: My experience is that different company cultures genuinely attract different people. The people we attract tend to be more idealistic and care more about the essence of technology — that's quite obvious.
LatePost: Any patterns? For example, which companies or industries do people come from that you find more reliable?
Qian Wang: New graduates. Because this industry really doesn't reward experience — almost no one has done this before, everyone is in the first cohort. Recently we're also seeing people come out of big tech or startups who have actually trained models — some from large model teams, some from autonomous driving. We prefer to hire people who previously worked on large models.
LatePost: Why can't autonomous driving companies do embodied intelligence?
Qian Wang: First, generally speaking, autonomous driving's understanding of large models is still somewhat lagging in certain aspects.
Second, autonomous driving and robotics aren't as matched as many people think — it's not 100% match. Autonomous driving has no physical contact; robotics involves a lot of contact. The technical core is different.
Third, autonomous driving has extremely high safety requirements, so people coming from that background tend to have somewhat inconsistent ways of thinking. The latter two points are secondary though — mainly it's the first point.
LatePost: Can't other large model companies do what you do?
Qian Wang: This isn't purely a large model problem. It also involves hardware, systems, the randomness of the physical world, plus experimentation and organizational management issues — fundamentally mismatched with large model team DNA.
Large model teams are like the air force. An excellent pilot plus a plane and you're flying. How you shoot down enemy aircraft depends on individual combat ability. At their core, large model companies are relatively loose labs composed of top smart people.
Hardware teams are like the navy. You're on a ship where every position is highly coordinated — from the front end directly dealing with hardware and data, to data processing, to model training. The chain is really long; one position fails and the whole ship sinks.
LatePost: How did you overcome this genetic conflict?
Qian Wang: Finding the right people. Also technically speaking, the action modality is different from language and vision — you need to develop a new set of methods to utilize action data. This itself has high technical barriers and genuinely requires a native embodied intelligence team to do these things.
LatePost: How well have your algorithm and hardware teams gelled?
Qian Wang: Basically no silos anymore — people can collaborate fairly well as an integrated whole.
LatePost: If you had to score it?
Qian Wang: 8 out of 10.
#1, No Bubble, Sector Shakeout
LatePost: Recently Omdia released a report: 13,000 humanoid robots shipped globally. The top companies were AgiBot, Unitree, UBTECH, etc. What do you think of this report, and what commercial progress will the robotics industry see in 2026?
Qian Wang: Robots still can't work yet. Commercialization is a bit like crying wolf. For the past two years everyone said it was the year of commercialization; now that it might actually be the year, people don't believe it. Because expectations have been透支 [overdrawn] too much — many people drew the commercialization pie chart too early.
LatePost: You think 2026 is the year of commercialization?
Qian Wang: Commercialization can begin. Not saying it'll be mature overnight, but at least this can start being done.
LatePost: How did you arrive at this judgment?
Qian Wang: Mainly that technology has reached a threshold. Reinforcement learning is now viable, and you can rapidly deploy on single-point products through few-shot learning. Foundation models need to be good enough for reinforcement learning to work — I think these are quite significant milestones.
LatePost: What's your commercial plan for 2026?
Qian Wang: Achieving positive ROI in at least some scenarios — that's the biggest milestone.
LatePost: Which scenarios? I saw you previously mentioned public services, elder care, and so on.
Qian Wang: Housework, cleaning, organizing — that's one category. Another is single-point vertical scenarios in industrial settings, like screwdriving. This is typical of things that previously could only be done by humans.
This year we'll see robotics commercialization landing, in a positive-ROI way. I'm quite confident.

LatePost: What's your view on the competitive landscape? Besides you, which other companies can achieve positive-ROI deployment?
Qian Wang: Most are probably overseas companies. For example, 1X has already sold several hundred units. Figure also has some industrial scenarios that are starting to work, close to being there — these companies are all quite strong.
LatePost: What about domestically?
Qian Wang: Right now in China, most companies are focused on singing and dancing robots — it's a different scene from overseas.
LatePost: How do you view competition with domestic peers?
Qian Wang: First, we probably need to define what "peer" means. Within the broader embodied AI category, there's one group focused on locomotion. That doesn't necessarily require AI at all — it's purely a control theory problem. Boston Dynamics started this way; they didn't use a single line of AI code.
These companies are actually operating on manufacturing logic — making better products, making them cheaper. That's fine, but it's a different direction from ours.
So we're on the AI side, some robot companies are on the other side. Of course we'll all move toward the middle eventually, but I think hardware is easy for us, while AI is hard for them.
There's another category of companies that mainly integrate resources — in some sense, they're more like real estate companies.
LatePost: What does the competitive landscape look like for each type?
Qian Wang: The singing-and-dancing robot category will probably cool down, and the field will see some consolidation.
We're starting to see that trend on our side too. By 2026, whether it's commercialization or models, you have to show something real. In 2025 we still saw lots of new entrants, but in recent months, basically no new players have entered the model or full-system space — because the elimination round has begun.
Overall, things will still improve, because robots are actually being deployed. Once the market scales up, everyone will see it's not just hype. If you can't produce something genuinely useful for years, you'll soon face a major trough like autonomous driving once did. I don't think robots will have that kind of trough, because they're already landing in real applications.
LatePost: Many people say embodied AI is overheated, that there's a bubble.
Qian Wang: I don't think there's any bubble at all. Compared to autonomous driving, compared to all previous major tech waves, embodied AI is far too small a sector in terms of resource investment, valuations, and funding amounts — not to mention you're another order of magnitude below the United States.
LatePost: Does the United States' funding advantage make you think you should have gone back?
Qian Wang: Long-term, China still has the bigger advantage. Whatever the industry, basically from 1 to 10, or 10 to 100, China always outperforms the United States. So if we can do no worse than the US in the 0 to 1 phase, or even do quite well, then long-term we definitely have the advantage.
LatePost: Comparing horizontally, you believe Independent Variable is doing the best, right?
Qian Wang: We do have some reputation technically within the industry. There really aren't many people today who truly understand how to build large models, especially almost none in the embodied space. Of all the embodied AI companies worldwide, we're the only one built around a large model team at its core. In terms of technical strength, we're definitely top-tier among startups.
LatePost: The biggest impression this entire interview leaves me is how confident you are.
Qian Wang: My judgments over the past two years have been fairly accurate. For example, we more or less actively gave up on commercialization over the past two years — looking back, that was the right call.
LatePost: I don't just mean these two years. It seems like you've been this way since your student days.
Qian Wang: That's what you'd call vision. I think my vision is pretty good.
LatePost: The way you approach things seems fundamentally different from most people.
Qian Wang: I figure if you're going to do something, do something that can be number one. If it were purely about making money, I might as well keep doing quantitative finance — no need to endure this much hardship.
LatePost: What do you usually do when you have time to rest?
Qian Wang: Sleep. I'm an introvert. Sleep is the priority. Wake up, read a book — that's nice.
LatePost: What's the most recent book you've read?
Qian Wang: Scientific American.
LatePost: Alright... I heard you also like browsing Bilibili. Any preferred content?
Qian Wang: No, just pure browsing. (At this, Wang recites the video titles currently on his Bilibili homepage: "Inside Google DeepMind's Lab"; "High Schoolers Dancing at New Year's Gala"; "The Rawest Cooked Meat in the World"; "Today's Dose of Joy"; "Airborne Wind Energy System Completes Grid Connection Test"...)


CES "New Species" Field Guide | Unity Ventures Portfolio @ CES 2026
