Dialogue | Qiming Venture Partners' Alex Zhou with Tashi Zhihang's Yilun Chen and Yuanli Lingji's Tang Wenbin: Macro Consensus and Micro Divergence on Embodied Artificial Intelligence

Continuing to explore general-purpose robots, the ultimate goal of AGI

The 2025 World Artificial Intelligence Conference (WAIC) "Qiming Venture Partners · Entrepreneurship and Investment Forum — Venture Capital Unleashing the Resonance Cycle of AI Technology and Applications," hosted by Qiming Venture Partners, was successfully held on July 28 at the Blue Hall of Shanghai World Expo Center.

During the panel discussion, Alex Zhou, Managing Partner at Qiming Venture Partners, served as moderator alongside Yilun Chen, Founder and CEO of Tashi Zhihang, and Wenbin Tang, Co-founder and CEO of Yuanli Lingji and Co-founder of Megvii, for a discussion on "The Singularity Moment of Embodied Artificial Intelligence."

Alex Zhou, Managing Partner at Qiming Venture Partners (left); Yilun Chen, Founder and CEO of Tashi Zhihang (center); and Wenbin Tang, Co-founder and CEO of Yuanli Lingji and Co-founder of Megvii (right)

Chen stated: "Embodied AI is currently the hottest subfield in the AI market. Embodied technology is advancing at an exponential pace, and we are already standing at the early window of the singularity's arrival." He identified four major trends in embodied AI technology: the gradual maturation of robot body control technology, the expansion of end-to-end technology from autonomous driving to robotics, the accumulation of data that is about to unleash Scaling Law effects, and the emergence of high-degree-of-freedom dexterous hand solutions. He also argued that embodied AI and autonomous driving share common origins in both task scenarios and underlying technology — model techniques can be reused, engineering capabilities can be transferred, and the experience and insights gained from the autonomous driving industry can aid exploration and implementation in embodied AI. Finally, regarding赛道 selection, Tashi Zhihang follows a "golden triangle" logic of high value, large scale, and high difficulty: targeting real needs that users deeply care about, addressing markets with substantial space where previous-generation robotics technology fell short, ultimately achieving the AGI end goal of general-purpose robots.

Tang shared core perspectives on technological development, entrepreneurial logic, and scenario implementation in embodied AI, demonstrating profound insight into this emerging赛道. He emphasized that his entrepreneurial初心 has always been robotics — from earlier entry through logistics robots to his current commitment to embodied AI, his greatest confidence stems from deep faith in technology, particularly the remarkable advances in large models, CoT, and Agent capabilities. Tang believes two necessary conditions for robots to evolve from specialized to general-purpose: first, precise perception of the physical world; second, planning and reasoning capabilities for complex tasks. He noted that whether robots can ultimately be deployed depends on two core factors: they must be usable and好用 — genuinely solving problems; and their economic model must work. These two conditions will most likely start from backend applications, then move toward commercial use, and finally reach consumer applications.

Selected highlights from the discussion:


01/

Embodied AI Technology Is Accelerating

Zhou: Thank you both for joining this forum. I still remember when Qiming Venture Partners invested in UBTECH back in 2015 — there weren't many investors paying attention to the broader robotics industry beyond humanoid robots and industrial robotic arms. I recall a founders' group chat that had just a few dozen geeks for a very long time. But starting two years ago, we counted over 100 companies in China working on embodied AI and general-purpose humanoid robots. Among all the AI subfields we're discussing at this forum, none has generated more startup activity than embodied AI, so everyone must be very interested in this conversation.

Please briefly introduce yourselves and your companies.

Chen: Hello everyone, I'm Yilun Chen, founder of Tashi Zhihang. Over the past decade, our team has been fortunate to participate in the development of some leading autonomous driving core technologies. As a sub-proposition of embodied AI, we went through the complete ten-year journey, from initial laboratory proof-of-concepts to today, where many of my friends can experience our products in daily life, and where it continuously improves everyone's travel experience every day.

In the coming decade, we hope to build more general-purpose robot forms and more powerful physical-world AI, accelerating the integration of these technologies into human production and life at greater scale and speed. We hope embodied AI technology can become an important engine for industrial upgrading in the next decade. Thank you.

Tang: Hello everyone, I'm Wenbin Tang. My first venture was Megvii, and today I'm representing Yuanli Lingji, a relatively new company focused on R&D and implementation in embodied AI.

We've been working on robots for a very long time. From Megvii's very first day, I wanted to first give robots eyes so they could see the world, but our entrepreneurial初心 has always been to build robots. Megvii initially entered through logistics scenarios, making our first attempt at robotics. Like my senior Chen, we've seen many technological variables this year that could enable the shift from specialized to general-purpose robots. We hope to truly use large models and robotics capabilities to bring the ultimate form of AI to the physical world — this is what we're working hard on now.**

Zhou: First question: as leaders in this industry, what major changes and progress have you seen in embodied AI, humanoid robots, and general-purpose robots over the past year? Can you share whether you feel more confident about this field's development?

Chen: I've personally always been very confident in this field. I think everyone can see at each year's WAIC that over the past two years, the pace of embodied AI and robotics advancement has exceeded all accumulated progress before that — this is quite inspiring. As practitioners, we anticipate that the development speed will only accelerate.

A year ago, virtually all robot exhibitions at WAIC were static displays. Now, in terms of whole-body domain control — locomotion and WBC — I believe this area is approaching a converged form. Another important AI development: end-to-end technology. Perhaps a year or two ago, academia had fairly strong confidence, while industry people still had doubts. But at least now, in the mobile domain of robotics, and in its largest scenario — autonomous driving — it has been fully productized, and people can experience its capabilities in daily life. In the manipulation domain, we've already seen enormous potential for leapfrog improvement at the laboratory product prototype level.

Third, I think something very important is that multimodal large models' fundamental capabilities have been continuously improving, and unlike pure language modality large models, multimodal models including vision and language still haven't hit a ceiling on Scaling Law — there's still enormous room for improvement. Combining these factors, I believe embodied AI is entering a period of accelerating progress.

Meanwhile, hardware technology is also maturing rapidly. We're seeing some very high-degree-of-freedom terminal forms, such as dexterous hands, with solutions approaching mass-production readiness. All this rapid development is quite inspiring.

Tang: I think fundamentally, the greatest confidence comes from large models reaching a certain critical threshold in CoT and Agent capabilities. I believe two necessary conditions for robots to truly become general-purpose:

First is precise perception of the physical world — something Megvii has been working on for many years. Whether from small models to large models, we've seen multimodal perception capabilities continuously strengthening, and they can now perform very well. Second is complex planning and reasoning capabilities.

Only when these two are combined can robots reach a general-purpose state. And today we see the development of Agents and CoT bringing us many surprises. So I believe combining these two points, from a technical judgment perspective, we're rapidly moving toward feasibility.


02/

Macro Consensus Gradually Forming,

Micro Diversity Still Apparent

Zhou: Very good. Regarding technology, we've covered a lot, but I'd like to explore further.

I recall when we invested in Megvii in 2014–2015, Qiming Venture Partners had its own investment thinking and logic. We believed that 2012's ImageNet was a turning point or breakthrough for deep learning, because after that, technology began converging — the industry's best and brightest were all striving toward one major direction. So we believed we could make bets on deep learning technology-driven companies like Megvii.

When we invested in Zhipu AI in 2022, and later in StepFun, we similarly believed that 2020's GPT-3 was the breakthrough point for large model technology. After that, technology relatively converged, with everyone working toward a common direction, which would surely yield good results.

When investing in Tashi Zhihang and Yuanli Lingji, we had much internal debate: has embodied AI technology converged? Or is it still in a state of百花齐放? If it's still百花齐放, the investor risk is substantial — the company we invest in today might have an excellent team, but if in three years technology converges away from that company's direction, wouldn't that be a major risk? Let's discuss: has embodied AI technology converged? Previously, large model development was constrained by data and compute. Does embodied AI face major bottlenecks preventing faster progress?

Tang: My judgment is that technology has not converged, because today whether in algorithmic frameworks, data sources, hardware forms and stability, or finally the sequence of scenario implementation — each remains an open question.

There is gradually forming consensus that technology is converging toward end-to-end, purely data-driven approaches, using VLA-like technical frameworks, and I think people also have some consensus about future technological development.

For example, on multimodality: people today generally agree that relying solely on visual guidance is difficult for achieving intelligence, because when humans interact with the physical world, we don't just perceive through our eyes — we use touch, and for things we can't see, we might probe with our heads. Can we learn from autonomous driving? How can we directly incorporate depth information into VLA? How can such multimodal data be fed into large models? I think consensus is gradually forming here.

But what does this model architecture look like? We actually don't know yet.

Some technical directions we're still exploring: today's VLA models are mostly single-frame. If you use VLA to drive a robot to cook and add three spoonfuls of salt, it actually can't add three spoonfuls, because after adding the first spoonful, it quickly forgets whether it added salt — from a visual perspective, the state with salt added and without looks identical. Since this model currently lacks a memory mechanism, we could of course build an external rule-guided mechanism, but how to give the model a native memory mechanism? I believe this is also a very important question.

Third, an internal research question: many companies today, starting from Figure, are proposing大小脑 models, but I don't believe大小脑 models represent an ultimate state.

The大小脑 model is essentially an artificial division by frequency — the brain area thinks, the cerebellum area executes, and their output frequencies differ, so we artificially split them into two models.

But is such artificial division a good approach? Is it intelligent? Actually not, because when humans operate, we think then act, and after acting, the state changes, so we think again. So how can robots form a dynamic, flexible chain of thinking and decision-making like humans? It might still be based on a single model, becoming a dynamic frequency and flexible frequency output from that model — this may be another open question.

So to answer your earlier question, I believe today's model framework is far from converged, with many problems awaiting solution. But precisely because of these open questions, I believe this is what fills us with passion and imagination for the future.

Zhou: Qi Yin (editor's note: Chairman of Qianli Technology) said that when Megvii was founded in 2011, it was student entrepreneurship during a wave of college student entrepreneurship, when the most common phrase was "jump off the cliff first, then assemble the plane while falling." But today's summary is: without first thinking through a complete technical and commercial闭环, such entrepreneurship would likely struggle to succeed.

This question is somewhat challenging — you just said there are still so many uncertainties, technology hasn't fully converged. Isn't choosing to start an embodied intelligence robotics company today essentially jumping off a cliff and assembling the plane?

Tang: I believe this is a question of the dialectical unity between "technical faith and pragmatic value." Because whether we're working on large models or going back to deep learning, without technical faith, no technology can give you certainty on the day it's born. If it already had very clear certainty, then the matter would already be finished, leaving no opportunity for startups.

So I believe it's precisely this uncertainty and technical faith that creates opportunities for startups. Therefore, I believe it's extremely important that within the team, people truly believe in this, holding genuine passion and faith in the technology.

Second, this process isn't just about faith — you need to be able to follow the true mountain-climbing path, finding camps for resupply along the way, with阶段性 commercialization that can produce results. So on this question, I both agree and disagree — it's a dialectically unified process.

Zhou: Please share your thoughts on this as well, Yilun.

Chen: I basically agree with Wenbin's perspective, but can interpret it from another angle. My view is that at the macro level, or on the long-term trajectory, I believe there is now strong consensus on embodied AI. But at the specific implementation level, each company has its own diversified thinking. **Let me share why I believe strong macro-level consensus is so important.

I previously went through a ten-year autonomous driving cycle where, at the macro level, there was prolonged strong non-consensus. For example: should robotic modules require decision-making and planning to use AI? Should they be handled separately from perception? Should everyone use maps? These were all matters of non-consensus that were debated for a very long time — at the macro level.

Now for embodied AI, at the macro long-term level, people's understanding is actually very unified. For example, we all agree data is very important. We all agree the final deployment form of this model will likely be end-to-end, multimodal, where vision and other sensors both play very important roles. We agree that imitation learning alone is probably insufficient — reinforcement learning is needed, perhaps even world model加持. On these points, there is common ground.**

But from a practical implementation perspective, the differences are substantial. **For example, on data: some believe many robots need to be deployed to gather substantial manipulation data; some believe large amounts of data need to be generated through simulation; some believe real-machine data is more important and should be collected "多快好省" through better methods. More specifically, regarding VLA which was just mentioned — I very much agree. I believe VLA represents three modalities: V for perception, L for language, and A for action output. So VLA defines the task input and output of the network, but what architecture should be designed in the middle? Does it need one network going straight through from end to end? Or are there some latent variable layers in between? Is imitation learning sufficient? Should reinforcement learning be adopted? What kind of reinforcement learning? Is world model加持 needed? These are what everyone is continuously exploring.

Actually, it's not just at the algorithmic level — it's the same at the hardware level, operating in a state of macro consensus with micro non-consensus.

For example, current general-purpose robot forms are basically two categories: bipedal and wheeled, representing different application domain trends. But even for bipedal robots, there are direct-drive joints and more complex transmission mechanisms that enable more balanced design between motors and their transmission systems — these all exist.

But I believe a macro consensus plus micro diversification is relatively healthy for this industry. It means everyone can rapidly iterate in a basically determined direction, defining their own认知 against each other, which will allow the industry to progress faster.


03

Past Industry Knowledge and Experience Can Be Highly Reused

Zhou: You previously led Huawei's first-generation fully ground-up intelligent driving technology self-research system, and shaped Huawei intelligent driving's global position today. You mentioned认知 — what认知 can be shared between the intelligent driving domain and today's embodied AI domain?

Chen: I think this is an excellent question. First, **autonomous driving technology and robotics technology have been同源 from the beginning. In fact, for a long time, autonomous driving's core technology mainly originated from two American robotics laboratories: Sebastian Thrun's laboratory at Stanford University (author of Probabilistic Robotics) and Red Whittaker's laboratory at Carnegie Mellon (lunar exploration robots). Through the DARPA Challenge, these converged into Waymo's main solution, which continues to this day. After 2018, autonomous driving technology began large-scale AI-ization, transforming from traditional robotics algorithm stacks with module-by-module AI-ization, to分层 end-to-end, to彻底 end-to-end AI-ization, making autonomous driving the first large-scale commercial embodied intelligence system.

I understand the reusability of autonomous driving experience, including technical experience, for the robotics domain in three aspects:

**First, direct technical reuse. Because robots, like cars, are also excellent embodied platforms for autonomous driving — they themselves require mobility capability, and this mobility capability is crucial for overall robot applications. Considering some commercial robot systems commonly seen today, their mobility technology is more similar to household robot vacuum technology. I believe directly upgrading from these technologies to more modern end-to-end technologies is very important for both application value and technical value.

**Second, some cognitive-level assistance. Autonomous driving has had enormous industry investment over the years, and one thing learned "the hard way" is that all AI in autonomous driving must be defined in time and space, not in two-dimensional images — this is extremely important.

In autonomous driving there's a well-known term, BEV — essentially a spatiotemporal concept. Defining things in spatiotemporal concepts has many benefits: regardless of any modality's input and output, they align on these very fundamental physical quantities of time and space.

From this perspective, my team prefers to call embodied AI "physical-world AI." As we heard earlier regarding pharmaceutical挖掘, that may be the chemical or biological world, but embodied AI inherently exists in a physical world, processing basic variables of time, space, and force. We believe a key factor for embodied AI's rapid advancement may be this认知.

Additionally, as the first large-scale deployed embodied intelligence system, autonomous driving has been through massive data冲刷, so there's clearer认知 of each method's capability boundaries — for example, the capability boundaries of imitation learning, of reinforcement learning.

**Third, direct engineering capability transfer. Robot hardware systems and autonomous driving hardware systems are basically completely similar in design, or in some foundational software systems — from chips,底层 software to communication middleware, they're basically highly convergent. And regarding the fast-slow dual system Wenbin mentioned, I personally very much agree with Wenbin's view: the fast-slow dual system is not the endgame, but it is a pragmatic consideration given current chips' memory wall limitations. So the asynchronous deployment of fast-slow dual systems, plus the two most important things for AI enterprises: one is the data pipeline, the other is training infrastructure — these can all be highly reused.

Zhou: I'd like Wenbin to address this as well. You built large-scale logistics robot deployment at Megvii. Comparing that experience to today's new-generation robot R&D, what do you think can be transferred over?

Tang: When we developed logistics robots back then, frankly, we were more seeking a focal point between market demand and technical feasibility. And the logistics industry is a very typical scenario — on one hand it could carry and validate our technology, on the other it had sufficient scale and clear demand.

As mentioned earlier, when Megvii was founded we wanted to build robots. At the start of our entrepreneurship, we began with eyes first, hoping to eventually have hands and legs to truly influence the physical world. We actually looked at many scenarios and found that logistics had several advantages: to some extent it was standardized. For example, the shipping container is one of the greatest inventions in logistics history because it encapsulates and standardizes many things, and this standardization makes automation and robotics feasible.

Logistics is actually an excellent scenario for robots to play a role — it has very large market demand, **with tens of millions of people working in warehouses globally, so demand is enormous. And because of its standardization, technology becomes feasible — this was the first very attractive point about logistics scenarios.

Second, in the process of building logistics robots, we actually paid substantial tuition, or learned a lot. One thing about building robots: we found many process links are embedded type, with predecessor and successor processes in physical space. In such process links, exception闭环 is extremely important. For example, in the digital world, if a virtual Agent or an App encounters an exception, you can restart the app and try again. But in the physical world, this can't be done — once a product is taken out, and a robot is transporting this product when our program malfunctions, how do you recover state? Its exception can't be solved by programmers capturing it, so we must design exception闭环 for the entire process. When you encounter this problem, how can you handle it to ensure production links proceed smoothly and completely? The actual cost of this may be much greater than imagined — this is a huge chasm from POC to actual application. This was the first thing we learned from logistics robots.

Today everyone sees many robot configurations, and internally we're also working on hardware forms. I think we learned another thing from logistics robots — fast isn't necessarily "fast"; stable may actually be truly "fast." We procured many robots, but their MTBF (mean time between failures) hadn't reached the requirements for truly long-term stable operation in scenarios.

And in this situation, large-scale deployment could lead to运维 disaster, where technical immaturity is compensated by service — such service is very "consuming" for teams, with large numbers of technical personnel and algorithm engineers needing to go on-site for various运维 tasks. We went through this once.

Finally, returning to robotics: when deploying in scenarios, these issues equally require serious attention. So I'm very grateful for this logistics robot experience.


04

Backend Manufacturing Scenarios Most Likely to Achieve Scaled Deployment First

Zhou: Very good. Everyone must be very interested: this WAIC gathered 150 robots, which seems very lively, but most remain at the stage of stage demonstrations. From the perspective of industry leaders, what will be the first or first batch of real落地 scenarios?

Chen: Actually, I think many robot scenarios are good scenarios. Let me share Tashi Zhihang's methodology for selecting scenarios — basically three sentences:

1. High value.

2. Large scale.

3. High difficulty.

We believe these three are self-consistent.

Zhou: High value, large scale, high difficulty.

Chen: High value means users have rigid needs with clear pain points. We hope for a larger product space so we can gather excellent people to work on this. And high difficulty is basically logically闭环: high value and large scale likely go together, and if the previous generation of robots could still solve the problem, then this generation of robots wouldn't have an opportunity. Our focus is on solving technical problems that the previous generation of robots handled poorly. From an application space perspective, robot practitioners and users have already shifted interest from showing off technology to deep thinking about use value — I believe this is a very good thing.

Any domain that can achieve scaled deployment is a good domain, capable of triggering the market's "singularity."

Zhou: Can you give a specific落地 domain? Can you reveal it?

Chen: From my perspective, the first domain with rigid needs and clear落地 potential is definitely manufacturing, because this industry already has large numbers of robots, and its pain points are very clear.

Tang: We also have some thinking about scenario selection, with several criteria:

First is positive gradient on the technology development path. Should we today deeply enter a vertical scenario? Internally we believe no — we must stay on the correct path of technology development, because many aspects of today's technology haven't converged. If we prematurely固化 the technology form, locking it into one scenario, we're sacrificing generalization to some extent. This isn't what we want to do, so we very much insist on advancing with a single model along the positive gradient of technology development.

Second, we simultaneously consider technical feasibility, as Qi Yin said: assembling the plane while falling off the cliff — some planes can be assembled, some probably can't be assembled today. For embodied AI using end-to-end pure data-driven approaches, going straight to 100% is very difficult, so we'll likely gradually progress from 90% to 95% to 100%. So how to find scenarios with relatively higher fault tolerance and tolerance for operation time? We believe this is very important.

Third, as senior Chen said, it must be a large-scale, high-demand scenario.

Specifically, Zhifeng's final prediction in his speech was very accurate — we also believe it will be more backend-oriented scenarios, such as industrial and logistics scenarios, because they're larger in scale, more密集, with more labor, so they generate greater value. Ultimately, whether robots can actually be used depends on two core factors: usable and好用 is the first point, because they must genuinely solve problems; second is whether their economic model works. These two factors will most likely start from the backend, then move toward more commercial applications, and finally reach consumer applications.

Zhou: Special thanks, and looking forward to both of you achieving great things in embodied AI!

Source | IPO Zaozhidao


Past Reviews

Qiming Perspectives | Alex Zhou, Qiming Venture Partners: 2025 Will Be the Year of Comprehensive AI Application落地 Qiming Stars | Multiple Qiming Venture Partners Portfolio Companies Appear at 2025 World Artificial Intelligence Conference Qiming Headlines | 2025 WAIC "Qiming Venture Partners · Entrepreneurship and Investment Forum — Venture Capital Unleashing the Resonance Cycle of AI Technology and Applications" Successfully Held

Founded in 2006, Qiming Venture Partners currently manages 11 USD funds and 7 RMB funds, with total committed capital reaching $9.5 billion. Since its establishment, the firm has focused on investing in outstanding early and growth-stage companies in Technology and Healthcare Innovation.

To date, Qiming Venture Partners has invested in over 580 high-growth innovative enterprises, of which more than 210 have listed on the New York Stock Exchange, NASDAQ, Hong Kong Exchanges and Clearing Limited, Shanghai Stock Exchange, and Shenzhen Stock Exchange, or exited through M&A and other means. Over 80 portfolio companies have become recognized unicorns or super-unicorns.

Many Qiming Venture Partners portfolio companies have grown into the most influential companies in their respective fields, including Xiaomi (01810.HK), Meituan (03690.HK), Bilibili (NASDAQ: BILI, 09626.HK), Zhihu (NYSE: ZH, 02390.HK), Roborock (688169.SH), UBTECH (09880.HK), WeRide (NASDAQ: WRD), Insta360 (688775.SH), Gan & Lee Pharmaceuticals (603087.SH), Tigermed (300347.SZ, 03347.HK), Zai Lab (NASDAQ: ZLAB, 09688.HK), CanSino Biologics (688185.SH, 06185.HK), Schrödinger (NASDAQ: SDGR), MicroPort EP MedTech (688617.SH), Sanyou Medical (688085.SH), Amoy Diagnostics (300685.SZ), Berry Genomics (000710.SZ), GenScript ProBio (688520.SH), Yuanxin Technology, Insilico Medicine, MediLink Therapeutics, LaNova Medicines, Zhipu AI, StepFun, Biren Technology, and others.