MaHui Entrepreneurs · 2025 WRC Series | On Embodied Artificial Intelligence, What Do Three Founders Have to Say?
From August 8 to 12, the highly anticipated 2025 World Robot Conference (WRC) took place in Beijing's Yizhuang district. This year's theme was "Making Robots Smarter, Making Embodied Intelligence More Intelligent," drawing over 200 leading robotics companies from China and abroad to showcase more than 1,500 cutting-edge exhibits. The number of humanoid robot manufacturers present set a new global record for this category of exhibition.


From August 8 to 12, the highly anticipated 2025 World Robot Conference (WRC) took place in Beijing's Yizhuang district. This year's theme was "Making Robots Smarter, Making Embodied Intelligence More Intelligent," drawing over 200 leading robotics companies from home and abroad to showcase more than 1,500 cutting-edge exhibits — with humanoid robot manufacturers setting a new global record for such events.
During the conference, the Ministry of Industry and Information Technology's SME Development Promotion Center, together with Source Code Capital and others, hosted a forum on the morning of August 10 titled "Industry-Finance Linkage: Reconstructing the Value Density of Robotics from Lab to Industry." Three Source Code Capital portfolio company founders delivered keynote speeches: He Wang, founder & CTO of Galaxy Universal; Tianlan Shao, founder & CEO of Mech-Mind; and Yongkun Wang, founder & CEO of Standard Robots. They shared their latest explorations in robotics and their observations on China's robotics industry.
Below are excerpts from their presentations, edited for length:
1
Galaxy Universal
He Wang, Founder & CTO:
Massive Data Requirements Are the First Hurdle for Embodied AI

I'm very glad to share Galaxy Universal's progress on our closed-loop, synthetic data-driven approach to humanoid robot commercialization.
First, a brief introduction. Galaxy Universal was founded in May 2023, incubated by Peking University and BAAI with the university's formal approval for technology transfer. In just two years, we've grown into a unicorn valued at $1 billion. Our mission is to put general-purpose robots to work across every industry and in every household.
Why have we moved so fast? It comes down to our positioning: building humanoid robots with genuine productive capacity. Everyone knows the humanoid robot market is hot right now, but in the long run, robots need real productivity to help meet the country's future manufacturing and economic development needs.
A previous speaker asked whether having data solves all problems. That's hard to answer, but I do know this: the biggest problem in embodied AI right now is that we don't even have data. Without data, solving problems becomes much harder, if not impossible. That's why companies like Tesla and their domestic counterparts are frantically hiring people to collect data. But robot data collection is far more difficult — users drive their cars and data flows back automatically; robots don't work that way.
Right now, humanoid robots pay teleoperators to collect data this way. This teleoperation is more exhausting than doing the task yourself, and currently, hiring people to collect data means you don't have enough robots, money, or manpower. In fact, embodied foundation models need data volumes comparable to language and vision models — on the order of trillions of tokens. We're still three to four orders of magnitude away from existing real-world data. This is the first critical barrier that embodied AI must break through. Once we have enough data, whether we can train embodied AGI goes without saying.
How close are current models to AGI? If embodied AI today could reach the level of ChatGPT or DeepSeek, it would be enormously helpful in daily life. But we're far from that because we don't even have the data. From my time at Peking University's lab to my PhD at Stanford University, my core idea has been solving the embodied data famine and the high cost of real data. The approach is large-scale physics simulation and synthesis: starting with massive 3D digital assets, generating large-scale motion trajectories through synthesis or reinforcement learning, then using sim-to-real methods so the trained models work in the real world.
In short, the entire pipeline is now complete. Many robot skills are trained entirely on synthetic data — or 99% synthetic, 1% real. The data pipeline works by first building the robot's dynamics and kinematics models, then combining object assets with environmental assets like materials and lighting to create a physical environment. The synthetic data pipeline generates motions, a simulator verifies they're correct and functional, and finally consumer-grade GPUs render everything to produce massive amounts of synthetic data.
For example: this is our synthetic dataset of various grasps across millions of rooms — humanity's first billion-scale grasping dataset. Training on this data produced the world's first grasping foundation model, GraspVLA, which can successfully grasp objects under any lighting condition in any environment. All these lighting conditions and tabletop objects were unseen during training, fully demonstrating the generalization of embodied foundation models. Ask it to grab a shopping cart, an excavator, swimming goggles, or a voltage tester — one command and the model gets to work.

Compared to Tesla's teleoperation, which collected 100,000 data points for battery disassembly with 40 people over a month, our grasping foundation model needs just 200 data points and half a day for one person. Train on a Nongfu Spring bottle, test on Nongfu Spring and Oriental Leaf water containers — showing zero-shot generalization. The robot handles changes in bottle color, number in a row, height, and cap size, all different from the training example, yet still generalizes.
This shows the true value of embodied foundation models: the ability to generalize by analogy and emergent capability. The GraspVLA model made its debut at both this World Robot Conference and last month's World Artificial Intelligence Conference with fully real shelves where users could order various items — hanging packages, soft pouches, bottled yogurt, beverages, puffed snacks — all delivered to you. This is currently the first embodied model capable of handling dozens of items with different shapes and softness levels in one go.
Behind this, relying entirely on real data collection would be enormously expensive — the economics wouldn't work for users. Over 99% of the data volume comes from synthetic data. At this World Robot Conference, we also debuted a major upgrade: ambidextrous, dual-tasking capability. Users can order two items at once; as long as they're within reach, the robot grabs both simultaneously — not just bottles, but perhaps a bag in one hand and a bottle in the other.
Since WAIC, our GroceryVLA model has been upgraded again in just two weeks with simultaneous dual-hand picking. These capability upgrades all depend on our use of synthetic data, including synthetic human-following data — all synthetic, not real.

Models trained on this data can enable robot dogs to follow people through shopping malls. This was filmed live at Beijing's Wanda Plaza on Children's Day with no teleoperation — just real-time visual input and predicted motion trajectories, determining where to step next. You can see all kinds of clutter in the environment, changing lighting, and many people blocking the way.
Overall, our virtual-real fusion approach to data collection, originating from academic ideas, is now being applied in industry. One vivid example: we've launched the world's first 24-hour robot-operated smart retail warehouse. Today in Beijing, if you order medicine online, there's a good chance a Galaxy Universal humanoid robot is picking it up inside the pharmacy. We currently have 10 stores in Beijing in real operation. By year-end, we'll open 100 such pharmacies across Beijing, Shanghai, Shenzhen, and other cities.
So I've always believed this technical approach can truly enable humanoid robots to rapidly deploy across commerce, industry, healthcare, elder care, and medical applications.
2
Mech-Mind
Tianlan Shao, Founder & CEO:
The Real Key to Embodied AI Is Intelligence, Not Form

Today I'd like to share Mech-Mind's observations on the evolution of embodied AI technology and its commercial applications.
Regarding application scenarios, we roughly categorize them into three types. First, serving objects — typically manufacturing and logistics, where the end goal is to produce, process, or transport something. Second, directly serving people — restaurants, hospitals, hotels, homes, where the goal is direct individual service. Third, research, education, and entertainment, which sit between the other two, providing knowledge and emotional value.
Mech-Mind's approach moves from left to center, starting with manufacturing and logistics. Some companies move from right to center, with humanoid robot firms often beginning in research, education, and entertainment. Both left and right are somewhat easier than the "serving people" middle. Our scenario selection is based on several dimensions, primarily environmental controllability, willingness to pay, user professionalism, and task variety per position.

For directly serving people, environmental controllability is very low. Homes are the classic example — we can't even control every detail of our own homes. Coffee shops and hotels also have high randomness, and willingness to pay is relatively low compared to manufacturing.
Whether users are professionals is critical — we consider this the most severe challenge. Non-professional users mean very open-ended usage with many potential safety issues. In factories, robot products are operated by trained professionals with dedicated protocols. Once in malls and hotels, there will inevitably be children and people completely unfamiliar with technology interacting with robots, making safety issues very prominent. Willingness to pay can improve as AI sophistication increases and costs decrease, but the user professionalism problem is hard to solve — which is why directly serving people is so difficult.
On technical evolution, I particularly want to emphasize: artificial intelligence is the most critical part of embodied intelligent robots — it's the real intelligence. But there's false intelligence in reality, such as robots performing fixed movements. That's not real intelligence. Teleoperation isn't real intelligence either. These remain at L0 or L1 level intelligence — fixed actions or remote control, perhaps with minor adjustments like balance or variation.
Mech-Mind is the world's first company to achieve scaled application of L2-level intelligence. L2 is still single-task, but requires enormous autonomous decision-making, particularly in vision and motion — perceiving the environment, understanding it, and making corresponding decisions. L3's defining characteristic is multi-task capability.
In terms of intelligence evolution, my personal assessment is this: right now, L2 intelligence has made good progress. We've seen many applications and rapid growth in recent years, but penetration still has roughly 100x growth potential. So large-scale L2 growth in coming years is foreseeable, and L3 deployment is also foreseeable in the visible future — where we're making very good progress.
At the just-concluded WAIC, we demonstrated robots autonomously folding clothes, humanoid robots retrieving items from shelves, continuous demonstrations of dexterous hands manipulating various objects, transparent items, and diverse commands given through natural language — all live, continuous demonstrations.
In manufacturing and logistics scenarios, ten operations per minute is very common; in factories we achieve about one operation every 1.5 seconds, 40 per minute. A robot with a 0.1% error rate making a mistake every two to three hours is acceptable. A 1% error rate means a mistake every 5-10 minutes — not even enough time for an operator to eat. That's a huge difference, both a very low bar and a crucial one. Put simply, robot failure frequency needs to give operators time to eat. That's what we can currently achieve.
I believe the real key to embodied AI is intelligence, not form. Our largest equipment is industrial gantry systems thirty to forty meters long, weighing dozens of tons. Our smallest handles things at the micrometer scale. Form factor makes no difference to us — they're all operable robots in terms of data.
Today, many robots can already achieve long-duration continuous stable operation with reasonable customer ROI. Industrial robots have reached this point, collaborative robots have too, and mobile manipulators are rapidly approaching and in some cases exceeding it. Humanoid robots are still working toward it — they've made very clear progress, but need a bit more time.

Mech-Mind's robots now serve many scenarios, work with many equipment types, cover many industries, handle many processes, and manipulate many object types. Yet we have only one software platform — not a single line of custom code. Half our business comes from overseas with strong growth, annual deployments reaching tens of thousands of units, and the code behind these robots is identical. That's true generality.
Generality sometimes has its own falsehoods. It's easy to fantasize that humanoid robots can do everything, but they may be running specialized programs, only doing one thing — dancing, grabbing items, moving boxes. That's not real generality. I can proudly say we've truly achieved generality: the same cameras, the same software, working with diverse equipment.
Of course, I also believe generality has boundaries — not everything can be generalized. We consider eyes, brain, and hands to have relatively good generality. Even humans aren't general-purpose; physically we're all different — tall, short, fat, thin. Take automotive parts and tire assembly: objects vary enormously in size, dimensions, and weight, sometimes handling things tens of meters long, where mechanical components are hard to fully generalize.
Or moving steel plates: having eight humanoid robots circle around and lift together like people would be very inefficient. My view is that mechanical equipment is hard to generalize — mainstream industrial robot manufacturers typically have over 200 models. This reflects that industrial robots, combined with some forms and other forms, ultimately need at least hundreds of forms to cover mainstream applications.
True generality comes from several places. First, the "eyes" must be good. Good things are all alike; bad things are each bad in their own way. When perception is very accurate, with low noise and minimal deformation, it's good. Our products perform well on semi-transparent objects and under ambient light interference, providing excellent support for subsequent intelligent algorithm generalization.

Second is the "brain" — providing standardized, generalized software that adapts to complex environments across many scenarios. Finally, "hands" that can adapt to various robot forms and grasp common objects.
We do artificial intelligence, not imitation. I particularly want to share something: artificial intelligence is not magic. We cannot defy mathematical and physical laws. When your inputs have insufficient signal quality, excessive noise, or significant ambiguity, no matter how powerful the AI, it cannot achieve 99.99% accuracy. With sensors, using lower-quality ones can achieve certain results and make demos work. But when we want to achieve 0.1% error rates, or even 0.01%, across very diverse, various real-world conditions, sensor requirements become much higher.
So I've always been very grounded. Mech-Mind is truly a generalization company. The reason is simple: customer needs are destined to be fragmented and dispersed — this is robotics' fate. I want our products to be configurable, programmable, solving industry problems through intelligence — general-purpose equipment. We deliver the same software to all industries and all customers.
This is actually very difficult, because every deployment brings 10x challenges. Providing something customized versus providing something generalized that covers specific needs through general capability — the latter is 10x harder. But when every deployment is 10x harder, adding them all up and truly doing it well enables replication and scale.
Finally: although Mech-Mind has already achieved globalization, standardization, and generalization — MIR reports show we've been China's market share leader for five consecutive years, nearly 4x the second place, and also rank first in Europe, America, Japan, and Korea — we still believe robotics is an extremely systematic undertaking with a weakest-link effect. Everyone should first explore and break through in areas already capable of scaled deployment: industrial robots, collaborative robots, mobile manipulators. We're now also working with humanoid robots. But ultimately, every link must hold up; overall reliability and success rate depend on every component.
Standard Robots
Yongkun Wang, Founder & CEO:
The Main Bottleneck for Embodied Robots Is "Intelligent Software"

I graduated from Harbin Institute of Technology, majoring in automation. In 2015, I started my business right after completing my master's degree. From the beginning, I embedded my goal in the company name — Standard Robots — hoping to provide standardized robots for the industry.
We chose to enter the industrial sector based on two core considerations. First, the country had issued a series of top-level manufacturing plans, setting ten-year development goals for China's smart manufacturing industry, specifically mentioning the promotion of autonomous intelligent robots in production and manufacturing. Though that term has now become "embodied AI." We saw this enormous opportunity ten years ago and aimed for it.
Second, we wanted to build a technology-leading enterprise. Across industrial, commercial, and consumer scenarios, we believed industrial technology offered more long-term accumulation and growth potential, so we ultimately chose the industrial path. Upon entering this field, Standard Robots established a long-term mission and vision: to drive smart manufacturing transformation through automation, digitization, and intelligentization, pushing robots toward more general and more intelligent evolution, providing standardized smart manufacturing solutions for global industry.
Standard Robots has been rooted in the industrial sector from day one. We pursue starting from mobile robots, based on environmental understanding, adding different modules — from early lifting modules, conveyor modules, and towing modules, to now single-arm and dual-arm configurations — ultimately enabling increasingly rich functionality.
I believe China will be the first to birth embodied intelligent robots and general-purpose robots applied in factories. China has the world's richest industrial supply chain and application scenarios, which are the best training grounds for robots. Because having scenarios means having data, having data means having intelligence, and having data understanding and process understanding enables the evolution of generality and intelligence.
We truly focus on frontline manufacturing scenarios, creating incremental value for customers by replacing manual labor, rather than blindly pursuing unrealistic generality. We discover real problems in frontline scenarios, not imagining product features in labs. The difficulty here is that factory scenario access and industry know-how requirements are very high. From our vantage point in this industry, we see that many products are mostly imagined; once actually deployed in industrial scenarios, numerous long-tail problems emerge that only deep frontline engagement reveals.

Take the automotive industry. Previously, because cars had many configurations, buying a different model or color didn't mean the factory produced in order — they produced more of what was ordered more, so ordering early didn't guarantee earlier delivery. But now some factories can produce completely to order: order first, any configuration gets built. Because 100 cars on the factory floor may all have completely different configurations. This requires all information systems and execution units — from process coordination, logistics equipment coordination, robot coordination, to human coordination — to be highly customized and highly flexible.
Standard Robots' goal is to use embodied AI technology to create "Super Workers" in factories.
We're an early entrant in the industrial AMR space, leading in many core technologies. But customers don't care about technical leadership — they care about eight words: safety, reliability, intelligence, efficiency. These are prerequisites for effective technology. Our approach: we achieve safety and reliability through self-developed underlying technology and continuous iteration on a single core technology platform; we achieve intelligence and efficiency through accumulating massive scenarios and rationally allocating tasks across multiple robots.
Currently, the main bottleneck in embodied robotics isn't hardware, but "intelligent software" that gives robots autonomous decision-making capability.
The essence of robot manipulation is "physical contact" with objects, and this behavior makes problem difficulty increase exponentially — completely different from autonomous driving. If robots could think through and execute tasks following complete step-by-step reasoning, results would likely be better. This is why we introduced a hierarchical architecture.
On technical path, we believe end-to-end VLA is a potentially huge future direction, but it's not currently the only way to solve industrial scenario problems. We adopt a hybrid architecture of brain-cerebellum layering, individual intelligence and swarm intelligence. This is a more pragmatic, safer, more deployable path that can quickly create value for customers.
For individual intelligence, we want one robot to gradually iterate from autonomous driving to mobile manipulation to embodied AI. We use fast-slow systems, through a large-brain small-brain layered architecture, combining general reasoning models with deterministic task execution to guarantee overall robot execution efficiency.
For swarm intelligence, real industrial scenarios never have just one robot. Even a fully evolved super robot in a factory needs instructions from information systems, not human commands, and must interact with other equipment and robots to complete tasks.
We achieve cross-form swarm coordination and customer system coordination through RoboVerse, rather than only emphasizing individual generality. This faster enables the complete commercial closed loop of "real scenario application — data flywheel — intelligence improvement — accelerated mass production," achieving intelligence growth through scaled commercial deployment. Currently, excessive emphasis on individual intelligence generality makes swarm intelligence critically important.
So the VLA path more aligned with industrial reality is a VLM-empowered + deterministic execution large-brain small-brain layered VLA architecture.

Humanoid robots are an important supplement to robot forms but not the only one. Manufacturing needs various forms of robots and specialized equipment to handle different scenarios and processes.
Standard Robots believes industrial scenarios will long have both specialized and general-purpose robots, coexisting across different manufacturers and brands. Therefore, unified protocols and software environments are needed to control and schedule them, enabling multiple robot types to collaborate in the same scenario and jointly complete tasks. We want all robots to speak "Mandarin" — this was the original purpose and meaning behind developing RoboVerse.
For China's manufacturing intelligence, I believe four steps are needed:
First, worker mechanization. Numerous robot companies, whether industrial or other types, are at the pyramid's base — this goal must be achieved by hundreds or thousands of companies together. Then management software-ization. Within a scenario, software is needed to coordinate all types of robot equipment, information, and so on. Then factory digitization. With software comes data structure, with data structure comes data沉淀, and with data comes final data intelligence. Finally, only on the foundation of the first three can manufacturing intelligence be achieved.
These four steps form a pyramid, layer by layer upward. Many companies work from top down, from intelligence to worker mechanization — this cannot succeed. I believe it must be bottom-up; without the foundation below, manufacturing intelligence is a castle in the air.

Standard Robots' goal is to focus on the industrial sector, and through four five-year development plans over twenty years, become a true China smart manufacturing leader starting from underlying robot technology,打通 the complete bottom-to-top technology chain, evolving from China's Standard Robots to the world's STANDARD.





