Zhang Wei of LimX Dynamics: Humanoid Robots Are Essentially an AI Application | Agentic Era

"Serving People, Not Process"

The true starting point of the Agentic Era isn't just the leap in model parameters, nor is it merely breakthroughs in VLA technology. It's intelligence genuinely entering the physical world — perceiving, deciding, and bearing the consequences of its own actions.

Humanoid robots are among the clearest manifestations of this era.

At this year's AGM, we invited Professor Zhang Wei, founder of LimX Dynamics and an Oasis portfolio company, to share his insights live. His take was direct: humanoid robots aren't just "machines with two legs"; they're the most important AI application scenario for the future. They shouldn't be treated as "specialized machines" — they should become the next generation of general-purpose intelligent super-terminals, like the iPhone.

And the industry's acceleration — from cerebellar foundation models to closed-loop whole-brain systems, from engineering iteration to cost reduction — is turning past trend predictions into today's reality.

This article is adapted from his speech that day. Full text is approximately 8,500 words, reading time about 20 minutes.

Enjoy

Oasis Capital was our early investor, an important partner, and one of our few board members.

When Jinjian asked me to come share, honestly I was very nervous. The longer you've been in this industry, the less you find worth sharing. Because the industry is developing so fast, and investors hear roughly the same things at every conference. There are even terms I first heard from investors' mouths. I think everyone is pretty impressive; there's actually very little of genuine sharing value.

So today I'll share my views on this industry from my personal and company growth perspective, simplifying the technical discussion.

In 2022, our company was founded in Shenzhen.

Our team, myself especially, spent a long time early on doing research related to robot motion control theory.

After 2017, I focused on applied research related to humanoid robots, early on more theoretical, later more applied. Back then, doing robot experiments was extremely difficult. As Professor Gao Yang just mentioned, we had limited funding and could only collaborate with others.

The collaboration with Cassie at that time cost about $200,000-plus. In 2020, working with Digit also required $200,000-plus. If the robot broke, repairs in the US took two months locally, so doing one experiment was extremely luxurious.

Compare that to iteration efficiency back then — now it's probably ten times faster. Just getting a robot to walk was something to be proud of, and very difficult.

Our first website looked like this.

What satisfied our original aspiration was just this one image. We called it "Cross the Limit" — hoping to break through the boundaries of robot motion and capability, letting it truly enter our lives.

Then we had a slogan: "Redefine General Purpose Robot." Why "Redefine"? Because most people who came to interview didn't believe it. They said, what's humanoid? Why does it have to be humanoid? Too much trouble.

We had to explain, to Redefine — redefining the commercial value and technical evolution path of humanoid robots, letting them couple with each other.

Then at the end of 2022, we debated with Wang Yuquan, then the "anti-humanoid fighter" and founding partner of Haiyin Capital. I was pro, he was con. He said there's no need for two legs; I said humanoid robots are the general future. Every question and answer from that debate, placed here today word for word unchanged, still holds.

The debate still exists, just with slightly more people on the pro side now.

Looking back, I personally firmly believed this was useful. But my estimates of technical evolution and trends were still conservative. Many developments happened much faster than I imagined.

My conclusion at the time was: you can't use today's technology to predict an exponentially growing track. Because back then, only a few people at Boston Dynamics could get a humanoid robot to take a couple steps. There were no large models, no VLA. You couldn't imagine things could develop this way. Now, if I tell you, getting a humanoid robot to walk — a master's graduate might be able to do it in a month. Before, it took the world's top scientists working for ages to figure out, and it still wasn't particularly stable.

A lot has happened in these three years.

I call this Exponential Growth — looking locally it all seems linear, you can't feel it day by day, but when you step back, such huge progress. Bill Gates said a few days ago that he spent about thirty to forty years researching speech recognition, invested billions and countless people into it, and after large models emerged, it was solved in maybe half a year. Now many companies' one year of work might equal what Boston Dynamics did over the past decade. So everyone has already seen exponential growth in large models — OpenAI's usage has exceeded trillion tokens, Google's monthly tokens exceed 1.3 quadrillion, and Doubao already exceeds 500 billion daily.

But I think the essence of humanoid robots isn't robots; it's actually an AI application. Just as AI large models empowering law requires legal professional know-how, specialized knowledge, it needs a professional Agent. Humanoid robots are also a specific application — just one that I consider the most important and the most difficult, slightly behind now, but humanoid robots are on the eve of explosion.

In 2024, roughly 2,000 or a few thousand units were produced; for 2025, predictions now seem conservative, definitely exceeding 10,000 units, roughly 5 to 10 times growth. If you count smaller humanoids, it's even more. Previously Goldman Sachs predicted millions in three to five years, still a linear prediction. We think it's in the early stage of exponential development. If we use current technology, many things are unimaginable.

Optimistic people might be like this — because new technologies will emerge, things we thought would take a lot of time may instantly become very simple.

Robots are actually the best carrier of intelligence, the super-terminal of the future. Everyone probably more or less agrees; what they disagree on is: when exactly can it be done? What's its value? How can this actually land? Lots of these questions.

How to make useful humanoid robots — that's a good question. After starting a company, I discovered that many "how" questions that aren't figured out or done well, the reason is often not knowing "why," so I still have to talk about why we need humanoid robots.

There's also something we don't know: after the "how" is done, then what? How to commercialize, how to make money?

So why humanoid robots? Let's ask from first principles. The question "is it absolutely necessary" is actually wrong. We need to look at what we need robots to do for us.

Its essence is to replace some of our daily labor.

First, many of our tasks require two arms. You can't do it with one arm — all our daily activities, like I couldn't do this right now with one arm. If I aggregate all tasks, probably over 70% of things require two arms to complete. So two arms are needed.

Once that's settled, where do we put these two arms? On a platform, on a table, where? Essentially it needs to move — we must transport these hardworking hands to where they need to work. So mobility is needed.

For mobility there are two choices: put it on wheels or on legs? Wheels work, legs work too. The criticism of legs is that they're unnecessary — so many places are flat anyway, wheels suffice. Another point: what scenario absolutely requires humanoid form, absolutely requires two legs? Some places can land without it first.

The essence behind these concerns is that we sometimes couple something's value with our ability to implement it. "Unnecessary" doesn't mean it's useless; it means it's too difficult, not worth the trouble — I think this meaning is implicit.

If I told you that under AI's foundation, legs and wheels are equally simple, many people might accept it. This is actually a fact. Looking at it now, legs aren't difficult, while arms grasping things — what Professor Gao Yang and Professor Suho are most expert at — is still quite challenging.

So it's not about what scenario absolutely needs robots; this is essentially using "specialized machine" logic to evaluate general-purpose robots. What's a "specialized machine"? Single scenario, single task, efficiency-first — simpler, lower cost. General-purpose robots are essentially multi-scenario, multi-task, solving long-tail needs. The chores we do at home every day — they're not things that need to be done repeatedly with high efficiency. How many times do you do laundry a week? They're actually long-tail demands where efficiency requirements aren't that high. So if you ask "what situation absolutely requires two legs," the answer might be: no single task requires two legs.

But if you ask, if I want to use one form to solve over 70% of people's common tasks, then it must be humanoid, must come with two legs. A third leg wouldn't help.

So if a single form needs to support multiple functions, that form absolutely has to be humanoid. Then some people say: I want cost-effectiveness. Previously, the technology wasn't good enough and costs were too high. But now there's a new variable: humanoid leg technology has matured. Walking — sure, we're still using remote control now, but that won't be a problem in the future, and cost won't be either. The cost won't necessarily be that much more than wheels, especially when it's more general-purpose. Because once it reaches scale, the cost advantage could actually be more significant than a "dedicated machine."

We can see this with smartphone cameras — they may offer better value than buying a more expensive dedicated camera. Why? Because after going through extreme generalization and scaling, costs drop dramatically. So there could actually be a scenario where it's cheaper than wheels. Jensen Huang also said he only cares about three endpoints: one is general-purpose humanoids, one is aircraft, and one is cars. Only absolute scale produces absolute efficiency, and humanoid is the form that satisfies this.

So we believe the underlying logic is this: sure, it seems useless now — waving its arms around, taking a few steps. But when it connects to enough apps, there will be a critical point, and once you cross that threshold, what I call the "Star Absorbing Technique" kicks in — everything can be developed on top of it. A phone used only for calling — you don't need a smartphone for that. Just for email — you don't need one either. But when many functions accumulate, you find everyone developing on it, and its costs gradually come down. That's our judgment of humanoid robots' future value.

Since we believe it's important and valuable, how do we do it well?

Now we're getting into the technical side.

First point: the humanoid robot body itself is actually easy to build. This is counterintuitive — the humanoid robot system seems complex because we're in the startup phase, but it's extremely easy to build, easier than aircraft or lithography machines. So why aren't humanoid robots being used yet? The reason is insufficient AI capability. It's not that robots are hard to build — as Jin Jian said, the fundamental issue isn't manufacturing difficulty, it's fundamentally a lack of AI.

This AI is somewhat special. It's different from large model AI — it's physical AI, connecting to physical systems and reality. At the same time, it's different from other physical AI in being more vertical: it's humanoid physical AI, and the overall system is relatively complex.

But conceptually it's simple. What can the whole system do? Just one thing: real-time computation and execution of commands. What's the input? A task given by a human — "go get a glass of water," "go fold some clothes." Previously, such command tasks required discrete remote control; now, with language, machines can understand — that bridge has been built.

After receiving the command, it needs to observe its surroundings in real-time — visual information and force information — and observe its own body state in real-time. That's basically it. The entire humanoid robot, embodied intelligence system — it's just one block, with very clear inputs and outputs. Whether you think of it as 10 million lines of code, or a massive model plus 1 million lines of code — that's debatable. But you can imagine it as one block. So humanoid is actually quite simple: what I see and what I feel, then deciding what to do now — it's just that mapping.

So where are we now? There has been progress. What do all the robots look like now, including ours and others'?

They look like this: there are some basic motion small models — "dance for me," "take a few steps," "do a somersault" — then switching between these small models via VR or remote control. So who is completing the entire physical system right now?

It's humans.

Artificial intelligence — when the intelligence isn't enough, use humans. Right now, people are telling it what to do first, like changing channels, doing things in a limited way. If this were the only usage pattern, then I'd agree its usefulness is quite limited.

But this is just today's level of development. Things we saw a few years ago are gone today; looking from now to three years in the future, I absolutely believe it won't be limited to today's development.

The humanoid body itself isn't hard to build — it's already converging. What's left is technology convergence, engineering, then mass production, cost, and who you're building humanoids for — that's where product capability shows. Another aspect is motion control: although training has compressed from a week or two months down to a few days or even hours, fundamentally, you can only do what you've pre-programmed it to do. We believe there will eventually be a motion foundation model — a whole-body motion control foundation model — where you don't need much post-training, just some language or posture prompts to generate the motion you want. That's important, plus the VLA brain. What we're doing now is eliminating the remote control and the human part.

Currently, the main approach is still collecting data, leaning toward imitation learning. Future development will definitely be about improving efficiency, reliability, and success rates. Eventually there will be a foundation model plus specialized skill fine-tuning.

But these are just models. We still feel what's missing is the ability to combine these algorithmic capabilities into a complex system — what we call the Agent OS system, which is the direction we're working toward next. For the cerebellum's motion control foundation model, we're working on it — in the future, you won't need to pre-program motions, you just tell it and it can do it. The brain foundation model is VLA, and there will be many small models combined with existing tools to complete each app. You don't need to reinvent a GPS model — use existing tools.

I think to build a good humanoid robot, you need at least three types of capabilities.

First: complete machine design and mass manufacturing capability. Unless you're purely a software company — as Gao Yang said, hardware-software integration is needed, and I think this is even more necessary, especially since the cerebellum part is tightly coupled with hardware. Complete machine design, mass production, cost — these are all hardware-oriented capabilities.

Second: algorithm innovation capability. Algorithm innovation doesn't mean copying someone else's algorithm. The next wave after the first wave of humanoids can copy homework and compete on vertical commercial operation capability, but right now it's still a stage where everyone needs to work together. As suho said, you need both innovation and commercialization, so you have to start early. So algorithm innovation capability is needed — both the big brain and small brain need it. Something rarely mentioned now, I think, is the need for AI operating system capability. In recent years, we've been focused on accumulating capabilities in these areas — I can only say we're seeing initial results.

For the body itself, I think this depends on your strategic choice. Most critically, several aspects affect its difficulty. First is size — internationally, 1.5 meters and above is full-size, below 1 meter is small-size. Small-size may lean toward entertainment and companionship. Slightly taller is relatively harder — previously difficulty increased exponentially, but it's better now. Another is degrees of freedom — many robots, if they're just gesturing, don't need that many degrees of freedom. Our body has 31 degrees of freedom, one more than Tesla's robot. Only with this many degrees of freedom can it better adapt to our environment. These two factors together determine the overall implementation difficulty and complexity.

Another thing to compete on, as Zhenyun from Yixi Biology said: "Others can't undercut your selling price with their costs" — this is quite important. I didn't care much about it before, but now I think it's extremely important. Cost control capability is critically important. Our current cost is about 50% of competitors', or even lower. Over the past year, we've reduced costs by 80% while increasing degrees of freedom by 10% — this is something that requires extreme attention to detail. Without cost control capability, you're basically not using China's core capability, so this is something we emphasize.

This is the mass-production-ready full-size full-degree-of-freedom robot we released in July this year. Both arms have seven degrees of freedom. Key components are self-developed, so the starting price is 158,000 RMB — that's the $20,000 Elon Musk mentioned. Our selling price is $20,000, and we still make money. That's what capability looks like.

Because the degrees of freedom are relatively high, using human data to train control methods becomes simpler. Previously it might have been extremely complex; now you just record some data and feed it back — this is directly related to the configuration.

That was about the body itself. For the cerebellum, this is where we've accumulated more. We're focused not just on walking, but on whole-body manipulation. Right now we can only gesture and dance, but eventually it needs to work. We're working on this too, it's just not at a stage where it can be deployed yet. You can't rely on coding every day to achieve this — you need data-driven approaches, even pre-training data that can drive large amounts of applications — this is quite important.

What we're doing now is already quite smooth — foundation model plus data for some fine-tuning, and it can do many movements. You can imagine in the future, ideally you give one sentence and it completes all the movements.

Besides whole-body movements, it definitely needs to do other kinds of movements too.

For example, this is teleoperation — it would be better with VR controllers. It's vision-based: the legs generate movements autonomously, not human-guided, while the hands follow human movements. With a motion capture suit, control becomes more precise, allowing remote operation of dangerous actions. These are the cerebellum's basic capabilities. We feel we've found the switch — everyone can just start working on it now.

One area that hasn't gotten much industry attention yet is the connection between the big brain and the small brain.

At the end of 2023, about a year ago, we had a video of a robot climbing stairs with real-time perception and full big-brain-to-small-brain integration. We've been investing heavily in this recently — getting the robot to see steps and using our TRON series for foundational algorithm validation. So it's capable of perceiving its environment, reacting, and then generating movement, rather than just moving blindly. This logic of connecting the big brain and small brain matters. The stair dimensions here are different, and it can still perceive them. Hopefully by this time next year, everyone will be able to see its more autonomous capabilities directly.

Right now the main criticism of humanoids is still the remote control. Honestly, I think people are overreacting to that — removing the remote isn't particularly hard. What's left is the big brain's foundation model.

This is where you need a foundation model that can perceive the environment and generalize, to guide the small brain's movements.

On the data pyramid and imitation learning that everyone emphasizes, I want to add a few points. First, large models are still in the pre-training phase. To actually deploy them, you still need post-training — you need to take that next step on real hardware, especially with reinforcement learning.

I believe intelligence fundamentally comes from data. Where does the intelligence in data come from? People. People are the ones who tell you this is a cat, this is a dog — the labeled data from before. In the process of transferring intelligence from humans to algorithms, there are efficient and inefficient approaches. The inefficient way is rote imitation, leaning toward imitation learning. The other is reinforcement learning, which is fundamentally about expressing intelligence. You're not telling it exactly what to do — you're just scoring it afterward, saying that wasn't done well, and letting it figure out the specifics. The data efficiency may not be high, but the process of generating data requires far less human intervention.

I think reinforcement learning is ultimately necessary for deployment, and that's a direction we're focusing on heavily. Whether it's pre-training to get started or reinforcement learning providing feedback to improve results, I call this "Divide and Conquer" — this combination: taking existing tools and current data-driven models and combining them to accomplish more complex tasks.

In the imitation learning process, we emphasize simulation as Professor Su mentioned, and we've also learned a lot from Professor Gao Yang. These three pieces are all critical for future deployment in our view. Within imitation learning, during pre-training, how do you improve data efficiency exponentially — that's one aspect. On the other hand, during post-training, how do you achieve reinforcement learning on real robot hardware.

Earlier this year, we had colleagues take cameras and have robots do tasks. After some time, we could use prompts to complete the same tasks in specific scenarios, across different robots. We call this "Data Recipe" — we're not all-in on real hardware or all-in on simulation, but have the full technology stack. When deploying in different scenarios, whichever you need, they can complement each other.

Reference article: "LimX Dynamics Releases LimX VGM Embodied Robot Manipulation Algorithm"

For the last mile of future deployment, we think skill composition is also critical. Robots can simply learn some Contact-Rich manipulation operations, then complete tasks in unseen environments. With just imitation learning, after collecting data, they can plan in completely unseen scenarios, relatively open situations — this matters too. Just like you don't need a large model to relearn the calculator or relearn weather forecasting, you just combine what large models excel at with existing knowledge and tools to get things done.

Finally, I want to share how we think about commercialization going forward. First, you have to look at what the industry landscape looks like and find your position.

In humanoid robots, there will definitely be core component suppliers — motors, reducers, robotic hands, including our joints. INXBOT is an excellent supplier in this space. There will also be technology solution providers, distinct from component suppliers — they might provide VLM models, data, teleoperation, even train certain deployment algorithms; these are data solution providers. Third, there will be OEMs in humanoids, just like in the phone industry. Fourth, and I think quite important, are humanoid整机 brand manufacturers. Fifth — and this is where humanoids differ from other robots — specialized machines won't have an app ecosystem because they only do one thing. Humanoids, being general-purpose, will ultimately explode through ecosystem and various apps, attracting everything to the platform like a gravitational pull. So the fifth category is application developers; some developers might make more money than the hardware manufacturers themselves. Finally, there are sales and channel service providers.

As for who's actually making the most money right now, it's not entirely clear yet, but the center of gravity is still around humanoid整机 manufacturers. Upstream suppliers revolve around them, and applications are built on top of them. Mainly because this is the hardest part — you need manufacturing, design, algorithms and AI systems, and product capabilities. If you break it down carefully, there are so many capability modules to build, and that construction process takes considerable time.

I mentioned the US-China dynamic earlier — I believe China will dominate this market at a faster pace, there's no doubt about that. Our position is整机 manufacturer, a humanoid整机 manufacturer. Our positioning is "Serve People, Not Process." So at the end of last year I said, our humanoid robots won't go into factories. They said that's too absolute. I said no, no — right now there's no need to bother putting two-legged robots in factories. The future we see is serving people, being with people, serving you in that environment. It's not serving assembly lines or tools or engineering projects. That's not its purpose, or at least not ours. These capabilities were unimaginable before, and now they can be achieved through data-driven approaches, though reliability still needs some time to develop.

We feel the technological variables are constantly shifting. Figure and Tesla are taking the big closed-loop approach — they're not selling yet, going straight for home scenarios, and they'll only sell when there are enough functional apps. We're taking a staged commercialization approach, laying eggs along the way. We sell smaller complete units too; as soon as an important app emerges, we start iterating — that's our logic. Of course you could say their approach is beginning with the end in mind. Ours is too, but you have to think through what that endgame actually is. At the same time, we're more open than they are — it's not just us playing alone. We'd love for everyone to help develop apps; we provide underlying capabilities. For things that no one else is pushing forward, we'll step in ourselves.

Analogous to phones, the hardware is essentially the iPhone, our full-body motion control platform is iOS, and we have development tools so that like Xcode, we can support various industries in deploying applications. Some applications we have to build exceptionally well ourselves; others are developed by developers on top of our platform — that's our goal.

For our current products, following the "laying eggs along the way" logic, the lower body can be sold separately with its own apps; the upper body can be sold separately with its own apps; and the full body has its own trajectory too. The lower body has a TRON series, which isn't here this time — its foot end has a three-in-one modular design, which we think is currently the best, optimal entry-level biped platform for humanoids. It's quite capable and extensible. Adding the upper arms, it can also do some mobile manipulation research. In all-terrain mobility, we're gradually seeing some app trends emerging.

Bipedal is a form that will definitely exist, but it hasn't been widely deployed yet, so we'll make it the best. This video shows a dinosaur form built on our hardware platform, something our developers made themselves — many parks could do this kind of thing. Didn't Yixi Biology just talk about resurrecting dinosaurs? We hope that after resurrection, we can walk together. (Laughs)

For humanoids, right now there's still some demand for performances, guided tours, and such, but scenarios are still limited — mainly because capabilities aren't there yet. But don't view its future with a static eye. We need to look with future eyes, plus a little imagination.

Video sped up 2x for demonstration purposes

What's made us happy recently is that we're truly seeing the embryonic form of the next app that can actually work.

We've managed to achieve a relatively integrated connection between the "cerebrum" and "cerebellum" of humanoid robots — this is fully autonomous, what we call closed-loop Humanoid VLA validation. No remote control, no positioning or navigation systems whatsoever. It relies purely on active perception to coordinate the entire body. Picking objects up from the floor, collecting tennis balls — you can imagine it picking up socks, slippers, clothes. It can transform "Pick and Place" into "Sweep." That alone has real value. The physical body and lower-level motor capabilities — you could call it Harvest — can absorb advances from the VLA field like a "Great Cosmic Absorption Technique." The more progress VLA makes, the happier we get.

The platform and foundational framework are in place. Next comes the realization of individual apps. Hopefully by version 3.0, seeing humanoid robots will be completely unremarkable. I came here to give this talk today — maybe next time I just download an app and it hosts for me.

That's some of the future we see. Very glad to share it with you all. Welcome to continue the conversation.