The Highs and Lows of Humanoid Robotics: Unity Ventures' Xiao Wang and Independent Variable Robotics' Qian Wang in Conversation

Patience is as precious as gold.

This year, embodied intelligence has reached a moment where consensus and divergence coexist.

On one side of the coin are embodied intelligence startups securing funding round after round. On the other is the heated debate over commercialization. View embodied intelligence through different time horizons, and you'll make opposite choices.

At present, the hardware and software approaches for embodied intelligence — humanoid robots in particular — have yet to converge. But after years of exploration, the field has arrived at a technical intersection between models and physical bodies. Without substantial R&D investment now, there will be no mature embodied intelligence companies in the future.

Unity Ventures began laying groundwork in robotics early, investing in mobile collaborative industrial robots in 2019 and moving into embodied intelligence companies in 2023. Recently, Xiao Wang, founder of Unity Ventures, and Qian Wang, founder of portfolio company Independent Variable Robotics, joined Tencent Technology's livestream "The Path to Embodiment" to discuss how large language models are fundamentally transforming embodied intelligence — covering core bottlenecks, technical pathways, and real-world applications. Independent Variable Robotics announced this week that it has completed two consecutive funding rounds totaling several hundred million yuan, with Meituan as the sole investor in its Series A.

Key takeaways:

  • Current humanoid robot shipments remain low because these machines still can't perform genuinely valuable tasks, remaining largely at the "demonstration" stage. For robots to achieve "utility," the core lies in autonomous manipulation capability, reasoning capability, and the integration of both. Once intelligence breakthroughs occur, shipments will surge.

  • China has advantages in industrial foundation and engineering talent, with the potential to become a major robot-producing country. In the long run, robots may become the third major hardware category most intimately connected to human life, after smartphones and automobiles. What's needed next is "patience" — industrial chain maturity requires coordinated effort across multiple critical nodes.

  • Robot "walking" is more of a hardware problem, while "manipulation" and "thinking" are more AI problems. Large models now offer entirely new methodologies that can break through the long-standing barrier of robots being unable to operate autonomously. What's most needed now are model systems that can directly control robots and enable physical interaction.

  • Among two technical approaches for humanoid robots, expert models are better suited to vertical tasks, while unified models hold greater potential. If you rely on systematic enumeration, once situations multiply, rules begin interfering with each other and the system becomes unworkable. Choosing the difficult but correct path of general-purpose models offers a better chance of achieving genuine breakthroughs.

Here are the highlights from this episode.

Source: Tencent Technology, AI Future Compass

Authors: Xiaoyan, Motong

01

Science Fiction Becomes Reality: Is "Humanoid" the Optimal Solution?

How do you view the gap between humanoid robots in science fiction and reality? What will future development look like?

Xiao Wang: Humanoid robots are not only achieving human-like gait but are also making gradual progress in facial expressions and other technical directions. Take the TV series Westworld — while it contains many science fiction elements, some concepts are gradually becoming reality: lifelike appearance, thinking and manipulation capabilities, and the ability to perform diverse tasks.

I believe these are no longer distant fantasies but rapidly advancing realities. In five to ten years, we may see humanoid robots with nearly indistinguishable appearances, capable of emotional companionship and household chores. With the development of large models, robots' comprehensive capabilities continue to improve, and the companies we've invested in are pushing in this direction.

Qian Wang: Current development of humanoid robots focuses mainly on two aspects: first, appearance that's more human-like, including walking posture, skin, and face; second, making their manipulation and thinking capabilities more human-like and more useful.

We're currently more focused on the latter. Independent Variable can already complete complex operations like pulling zippers, organizing flexible objects, and folding clothes. Looking at the current model performance from Google and Physical Intelligence, embodied intelligence is at a stage comparable to natural language processing in early 2019 when GPT-2 was released. We're now in a phase similar to the transition from GPT-2 to GPT-3. Although hardware, sensors, and models still have limitations, the potential for technical breakthroughs is clear.

In terms of locomotion — gait control and balance — robots have already reached or even surpassed human levels. As for appearance aspects like skin and expressions, there are no theoretical obstacles in the technology itself; they simply require gradual engineering accumulation.

On manipulation capability, we're also improving robots' thinking ability for complex tasks. The multi-modal "chain of thought" we've built can already support robots in long-sequence complex reasoning.

I believe robots will make astonishing progress in capability over the next five years. Just as no one in 2019 anticipated that a product like ChatGPT would emerge by late 2022, we have full confidence in robot development. The real landing of embodied intelligence will happen within a foreseeable timeframe, possibly even exceeding current public imagination.

Will humanoid form become the future standard? Is it an inevitable result of technical development?

Qian Wang: Regarding humanoid robots, I believe bipedal walking and human-like appearance are technically feasible, but whether they're the optimal path remains worth discussing.

One experiment Independent Variable is conducting involves letting everyone turn their skills and craftsmanship into fine-tuned models, allowing robots to acquire specific skills as easily as downloading apps. This approach can break through the traditional problem of human skills being impossible to replicate and difficult to circulate. From this perspective, the ultimate goal of AI and robotics is not merely to imitate and reach human levels, but to substantially surpass human capabilities.

While humanoid robots have irreplaceable value at the emotional level — because humans naturally form emotional connections with human-like appearances — in the long term, more efficient and capable non-humanoid forms may emerge. Just as humans didn't achieve flight by imitating birds but invented airplanes, the future form of robots isn't necessarily limited to the human-imitation path.

Xiao Wang: When investing, we focus primarily on what problems robots can solve and what scenarios they apply to. Robotics is a diverse concept — for example, robotic arms and forklifts in factories also fall under robotics. Humanoid robots are just one form, including bipedal, wheeled-legged, and other variants. Whether to adopt a humanoid form depends on specific problems and scenario requirements, not simply pursuing human imitation.

02

Application Challenges: From Exhibition Prototypes to Household Assistants

How do we break the current predicament of "only demonstrable, hardly applicable"? How can we drive broader application of humanoid robots?

Xiao Wang: Robots already in extensive factory use, while not humanoid, are quite common — robotic arms, pipeline automation equipment, and so on. In commercial closed-loop scenarios like hotels, restaurants, and cleaning, service robots have also been widely deployed. If we moderately broaden the concept of "robot," we could say they've already achieved some penetration in production and daily life.

But from now to the future, achieving humanoid robots with "human-like thinking and manipulation capabilities" still has a long way to go. The core challenge isn't entirely in hardware but in "intelligence." Only when robots can understand tasks and complete complex actions like humans can they truly count as "robots." While walking technology has made major breakthroughs, thinking and manipulation remain incompletely realized.

This is also why humanoid robot shipments are currently low. They still can't complete genuinely valuable tasks, remaining largely at the "demonstration" stage. Once intelligence capabilities achieve breakthroughs, shipments will surge dramatically.

Current large language models can be used to understand instructions and transfer knowledge, but cannot directly solve robots' manipulation problems in the physical world. What we need is an end-to-end system combining language understanding with action execution — this is the true "breakthrough point" for robots.

This requires teams to simultaneously possess hardware, large model, data, and systems engineering capabilities — precisely the most difficult part of current robot R&D. Once breakthroughs occur at this critical node, the robotics industry will explode. The real core lies in "generalizability of intelligent systems," just as Android was to smartphones.

Qian Wang: I agree. Although companies like Boston Dynamics and ASIMO have researched "walking" for many years with great progress, there remain significant deficiencies in "hand manipulation" and "thinking." The fancy robot manipulation demos commonly seen in the past were mostly based on preset trajectories. Every motion was repeating a pre-programmed path, not autonomously completed by the robot. Even some robots capable of fine manipulation, performing better than humans, relied on human teleoperation behind the scenes.

In fact, it wasn't until around 2018 to 2020 that robots truly achieved relatively complete breakthroughs in "autonomous grasping" tasks for the first time. For decades, the market built robot hardware with execution capabilities far exceeding human hands, yet robots simply couldn't operate autonomously.

To summarize, robot "walking" is more of a hardware problem, while "manipulation" and "thinking" are more AI problems. Large models now offer entirely new methodologies that can break through the long-standing barrier of robots being unable to operate autonomously. However, language models cannot be directly applied — they can solve planning, reasoning, and long-sequence cognition but cannot directly interact with the physical world. Therefore, model systems that can directly control robots and enable physical interaction are still needed, whether in end-to-end form or other implementations.

Of course, robots also have emotional value and display value. But to achieve "utility," the core still lies in autonomous manipulation capability, thinking capability, and the combination of both.

The industry goal is to reach consumers. What's the capital deployment strategy?

Xiao Wang: Overall, the humanoid robot industry chain is quite long, encompassing chips, joints, control systems, and "brain" modules like Independent Variable Robotics, plus deep integration with different scenarios. Only when capital forms consensus and concentrates investment in one direction can the industry mature quickly.

With the development of large models, robots are gradually gaining thinking and manipulation capabilities, with significantly enhanced generality. Meanwhile, at the hardware level, the gradual maturation of bipedal walking and dexterous manipulation hands also provides foundation.

China has advantages in industrial foundation and engineering resources. I believe China has the potential to become a major robot-producing country. In the long term, this will become the third major hardware category most intimately connected to human life, after smartphones and automobiles.

True commercial landing still needs five years or even longer to form a product form with good cost-performance ratio, consumer acceptability, and practical functionality. Therefore, both society and capital should give the industry sufficient patience.

Industrial chain maturity requires breakthroughs across multiple critical nodes. This isn't a task any single company can complete independently, but requires coordinated effort from multiple entities in multiple directions.

What are the constraints on industrialization? What key links are still missing in the current industry chain?

Qian Wang: First, price is an extremely critical issue. It involves product input-output ratio and the product-market fit (PMF) point, and PMF point design is the most important element in commercialization.

People's expectations for an item are strongly correlated with its price. For example, consumers buying a robotic vacuum for a few hundred or few thousand yuan don't expect it to perform complex tasks — just clean the floor well. That's a clear PMF point.

If we want robots to do everything humans can do, even surpass certain human capabilities, then we need to be willing to pay higher prices. The question is whether we can find an appropriate commercial landing point between the two — making products both practical enough to meet needs and acceptable for mass adoption. This is an important topic for industrialization.

Another constraint is industry maturity. For example, despite years of development, dexterous hands remain in early stages. Currently, dexterous hands with high degrees of freedom and strong reliability still command high prices, constrained by production volume and upfront R&D investment. But in the long term, costs will certainly fall to reasonable ranges.

Additionally, the industry has yet to reach consensus on key technologies. For example, technical approaches for dexterous hands, such as haptic feedback, haven't converged. Key subsystems remain in technical exploration stages, so more time and patience are needed.

In the future, as the industry naturally matures and AI capabilities continue improving, we hope to find PMF points that match market demand, thereby achieving shipment volume increases and substantial cost reductions.

03

Intelligence Core and Hardware Support: Diverse Technical Pathways

Some technical approaches favor achieving all functions end-to-end through large models, while others support systems engineering approaches combining multiple small models or traditional algorithms to realize complex functions. How do you view these two different technical pathways?

Qian Wang: There are substantive divergences in current technical approaches. One path involves building multiple expert models to form a function set or "skill library." The other is what Independent Variable is doing — implementing all functions in a unified model, a general-purpose model, a generalist model. I believe expert models are more suitable for vertical tasks; but to achieve general capabilities, a completely unified model is needed. This is precisely the fundamental reason for currently pushing large language models and multi-modal models.

Expert models have capability ceilings, while unified models have greater potential to break through existing boundaries. Of course, which path to choose also depends on ultimate application direction. Over past decades, extensive systems engineering strategies have indeed achieved some results, but the gap with people's expectations remains huge. Therefore, I believe more effort should be directed toward the general-purpose model direction — this is the direction more likely to break through technical ceilings.

Xiao Wang: We want robots to possess generalization capabilities, able to handle various uncommon problems. If you rely on systematic enumeration, once situations multiply, rules begin interfering with each other and become difficult to operate. While partial functions can be achieved in limited contexts, the system becomes unsustainable at scale. Therefore, I believe this technical approach may be attempted in the short term but isn't viable long-term.

I lean toward using large models for end-to-end solutions. Because wherever human intervention enters the design, vulnerabilities may exist, and any additional algorithmic adjustments may bring new problems.

The technical difficulties of unified models lie in model construction, data processing, and algorithm optimization, while also needing to consider adaptation to real-world scenarios. These challenges are extremely severe, but precisely because of this, only by choosing this difficult but correct path can true breakthroughs be achieved. The direction is clear; the key lies in data scale, algorithm optimization, and timing — still in exploration stages.

New model architectures keep emerging, such as Figure's Helix. From a technical perspective, what are its characteristics?

Qian Wang: Independent Variable's model architecture is similar to π0 in overall direction — both are end-to-end, fully unified models. Although for quite some time, the end-to-end approach wasn't widely accepted. But because robot hand manipulation has its particularities, many manipulation tasks simply cannot be completed without end-to-end approaches. Once manipulation difficulty exceeds simple grasping, traditional hierarchical models struggle to perform. Currently, "fully end-to-end, integrated, general-purpose models" represent a major development direction for embodied intelligence. Independent Variable's research team is also on this path.

Meanwhile, Independent Variable's model also differs from π0 in some aspects. For example, in high-level thinking, planning, and reasoning, Physical Intelligence typically uses separate independent models. Because π0's architecture itself involves less of these aspects, although it has existing VLM models as foundational backbones, after action training its language and vision capabilities somewhat degrade, thus requiring additional models for high-level architecture.

Independent Variable's model contains a complete capability system: thinking, reasoning, and low-level action control are all integrated. Our self-developed model WALL-A is currently the world's largest-parameter embodied VLA model, substantially surpassing π0 in task difficulty, high-level semantic generalization, action generalization, and modality alignment.

Our approach is fundamentally superior because as task complexity increases, non-end-to-end models all face a fundamental problem — how to combine modules. Once errors occur in preceding processing, subsequent stages are severely affected. The nature of robot manipulation problems drives Independent Variable to choose the end-to-end large model path.

This technology has now gradually developed to a relatively mature level. Whether using simulation or end-to-end methods, both actually stem from the characteristics of manipulation tasks themselves. We established the end-to-end technical route early on, believing that minimizing human intervention is a long-term trend — indeed, humans themselves struggle to clearly explain their own cognitive processes.

The rise of large model methodologies represents a major innovation and fundamental change in methodology. Whether π0 or Independent Variable's model, I believe both are on the right path. Even if future technical breakthroughs emerge, they will likely remain within the current (end-to-end) framework, unlikely to return to past hierarchical architectures or to the old paradigm of "expert models" (one or several tasks per model).

From a computing perspective, is it necessary to develop hardware specifically for robots? Does this direction have significant industry importance?

Xiao Wang: The core of robots is computing, and it needs to support AI operation. Traditional CPU and GPU vendors remain core suppliers of robot computing power, but some new smaller players will also enter this field and conduct specialized development. We've already begun deploying and investing in chips for the robotics field. Overall, development in this area remains in early stages.

Qian Wang: From our current perspective, automotive chips serve robot edge inference computing needs very well. Although these chips were originally designed for autonomous driving, autonomous driving has partial overlap with embodied intelligence in computing requirements.

There are also some differences. Compared to autonomous driving chips, robot chips have less stringent physical requirements. For example, robot chips don't need to withstand extreme high or low temperatures like autonomous driving chips, so costs are relatively lower. But from a computing perspective, existing GPUs and edge inference chips already satisfy embodied intelligence needs quite well.

In the future, autonomous driving models may not require as much computing power as humanoid robots, but as robot computing demands increase, embodied intelligence will need more powerful chips for support.

04

The Future of Humanoid Robots: Differentiated Development in the AGI Era

Please share your thoughts on DeepSeek's impact — how do you view this change?

Qian Wang: DeepSeek has had profound impact on the broader environment. Previously many people believed original work emerged more from the United States. DeepSeek has greatly changed this prejudice, especially overseas — people are beginning to realize China's strength in AI.

It hasn't only changed perceptions of China but also driven societal recognition of this issue. Therefore, for Chinese companies like ours conducting frontier exploration from zero to one, DeepSeek undoubtedly provides a good example.

At the specific technical level, DeepSeek's achievements provide valuable reference for us. But DeepSeek focuses mainly on language models and reasoning models, while Independent Variable specializes in embodied intelligence models — the two differ greatly in problem nature.

Many might think that since both are large models, they might be very similar. But actually, characteristics of each domain lead to enormous differences in technical approaches and specific choices. For example, autonomous driving and robotics differ in many aspects. Problems faced in robot manipulation are almost all absent in autonomous driving; while safety concerns in autonomous driving aren't encountered in embodied intelligence, so their technical routes are completely different with almost no possibility of reuse.

The comparison between us and DeepSeek is similar. DeepSeek-R1 focuses more on long-horizon reasoning and long chain-of-thought. Independent Variable also does chain-of-thought, but more in multi-modal form — such as predicting states of certain actions, or quality of actions — and doesn't require particularly long thinking. DeepSeek's long chain-of-thought and reinforcement learning are more suited to its domain, but for Independent Variable, these don't have direct technical impact.

Of course, DeepSeek is also advancing multi-modal models, which serves as reference for us, including some reinforcement learning algorithms. But overall, what DeepSeek does and embodied intelligence belong to two major directions of AI.

Xiao Wang: Two or three years ago, I said Chinese models wouldn't lag behind American ones. With Chinese engineers' mathematical capabilities and diligence, our models could completely match or even surpass American ones. DeepSeek proves China can create models on par with or better than the United States, making us more confident.

DeepSeek is like the open-source Android system, lowering application development costs and barriers. Developers no longer need to rely on paid APIs but can directly use open-source models, making application development lower-cost and more flexible.

If Independent Variable Robotics can successfully launch a large model for the robotics field, the entire industry may experience an explosion, just like the explosion at the application layer. By lowering costs, the robotics industry's application layer will welcome a true inflection point.