Oasis Capital Dialogue with Professor Yin Peng: Convergence and Flywheel
Constant Change

What counts as a robot in the truest sense? Can a data flywheel break through the "window paper" of robotic systems? Today we have with us Yin Peng, Assistant Professor at City University of Hong Kong and founder of MetaSLAM, to discuss the future of general-purpose robotic systems. Enjoy.

Oasis Capital: Please briefly introduce your research direction to our readers.
Professor Yin: Computers have operating systems, phones have operating systems — what we work on is robot operating systems. Robots need to interact with the physical world, so our focus is on how to perfect their operating system. This isn't exactly a new story. For years, whether in driving or warehouse logistics, all kinds of robots have required bottom-up holistic strategies. Over the past decade, we've concentrated on robotics technologies closely tied to people's daily lives — from localization, mapping, planning, and decision-making to perception — unifying them under a single architecture. It's somewhat analogous to autonomous driving, except robotics is harder and more sensitive. Current robots are specialized machines for specialized scenarios. In the long run, we hope to eventually achieve general-purpose robotic systems.
Oasis Capital: There are many terms for "robot" nowadays — Robot, Humanoid, and so on. If we're being precise, how do you define a true robot?
Professor Yin: Tracing back to the origins, "Robot" in science fiction referred to labor force — a transliteration of the word for slave. Robot was defined from the value and meaning of the machine itself: a tool. Humanoid came later, around the 1960s or 1970s, when people felt robots should develop in a human direction, giving rise to human-like, anthropomorphic robots. Going from the original definition, a true robot — without the "human" element — is simply a tool. If we want to赋予 human attributes or value, from a biological perspective, it should possess some form of self-awareness, spontaneous understanding of its surroundings, and a curiosity mechanism for the unknown that drives it to explore and optimize. It's not a living organism but a vessel. What distinguishes it from conventional things like airplanes or trains is its capacity for "reflection" — some attribute that makes people sense intelligence in it. Only with intelligence can it interact with humans at a deeper level, just as ChatGPT's large model interacts with people, leading them to perceive or feel that the model possesses logical reasoning and thinking-oriented attributes, capable of higher-level, human-like tasks. Specifically, there could be tool-type robots that replace humans in rubble or extreme environments; there could be companion or caregiving robots with emotional output, no longer pure labor. Of course, when it comes to ethics, whether such an emotionally endowed vessel counts as living or non-living — that's another question entirely.
Oasis Capital: By this definition, without the profound impact of ChatGPT, we probably couldn't even discuss achieving the robot you define?
Professor Yin: AI robotics isn't the first wave — it started back in the 1950s. This time, the fundamental change came from ChatGPT's rise, making people feel that models genuinely possess thinking ability and logical reasoning, completely different from the previous rule-based, human-defined if-else attributes.
Oasis Capital: Are there any robots today that match your definition?
Professor Yin: If we don't consider the hardware载体 and only look at logical thinking ability, ChatGPT 3.5 or 4 could be considered life forms with the intelligence level of a 10-year-old child. In the future, if ChatGPT 5 emerges, it might upgrade to an organism with college-level thinking logic. Without considering physical hardware entities, these already belong to robotic systems. If we must consider robots that interact with the physical world, there aren't many visible examples yet, aside from perhaps Elon Musk's humanoid robot or the 1X robotics company acquired by OpenAI, which adds physical peripherals onto the intelligent agent based on ChatGPT, forming a robot system with an本体.
Oasis Capital: Stanford's Fei-Fei Li proposed embodied intelligence VoxPoser, and Google also came in with PaLM-E. How do you view these different paths in this wave of development?
Professor Yin: This is quite an interesting point. Whether it's Google's PaLM-E or Fei-Fei Li's VoxPoser, they've all opened up a channel. Early ChatGPT could only be said to possess a certain degree of generalization attribute from NLP. Setting aside physical world mapping, ChatGPT had already emerged with basic logical capabilities at the text level.
But in real-world environments, before PaLM-E and VoxPoser, we basically hadn't seen any work achieve this degree — fusing real-world visual information, localization information, mapping information, sound information, even tactile information into a unified world model. I have to say PaLM-E and VoxPoser pioneered this field. Because the real world is so complex, we can't say they've achieved completeness, but they did very early on map the real world into latent space, allowing robots or ChatGPT's large model to understand the attributes of the physical world. That's what they accomplished.
Of course, this itself has shortcomings. Unlike NLP, which developed for so many years before ChatGPT emerged. BERT large models also evolved through many generations before deriving the currently viable paradigm.
The complexity of the real world lies in the cross-coupling of multiple information streams. For a robot to explore a room like a human, with spatial, sound, lighting information all mixed together, how do you map information from 3D or even N-dimensional worlds into a single dimension? That's the hardest problem. I have to say, Google and Fei-Fei Li's work represents excellent attempts, but there's still substantial work needed before reaching the kind of emergence we saw with ChatGPT.
Oasis Capital: At this point in time, if general-purpose robots aren't yet that "general," are there any practical落地 scenarios? Or only laboratory ones?
Professor Yin: From a practical落地 perspective, the combination of robots and large models, or general-purpose robot technology, is still in its early stages — perhaps requiring 3 years, 5 years, or even longer to develop. Looking back at NLP a decade ago, at best it could achieve the level of "Xiao Du Xiao Du" or Xiaomi's robot; even so, NLP had already been developing for many years. Robotics and CV fields both benefited from the rise of the NLP industry. Objectively speaking, if ChatGPT hadn't been born last December, most of the NLP we know would likely still be stuck in conventional thinking. But due to the sudden emergence of emergent abilities, demonstrating quantitative change leading to qualitative change, the development path for robotics is the same.
Currently, while there are no general-purpose robotic systems yet, whether it's扫地 robots, cleaning robots, or logistics robots, there are various system models and generalizations — no different from NLP's situation back then. NLP also had various subtasks for text processing. After ChatGPT emerged, many of these subtasks were eliminated. Similarly, once a general-purpose robotic system model emerges, conventional domain-specific robot generalization attributes will vanish completely. Before ChatGPT appeared, having such "pipe dreams" wasn't very realistic, but ChatGPT's emergence at least proved one point: in pure NLP, large models do possess high-level logical reasoning ability. As long as we can bridge the physical world and text world, robots can in principle also possess large model capabilities. How long this will take is indeed uncontrollable. But once it appears, the entire industry will be completely reshuffled.
We can also analyze OpenAI's journey — early on they weren't favored, market response was mediocre, and without ChatGPT at the end of last year, the company was at a "life or death" moment.
Era-defining achievements require a group of excellent scientists, investors, and major corporate support, accumulated over the long term, to present a "phenomenon-level" event. The robotics industry is relatively fortunate because OpenAI's trigger has let everyone see hope. With people and companies led by Elon Musk or OpenAI going all-out to冲击 robotic general intelligence, once systematic training proceeds规律ally, ultimately presenting general value, the societal attributes brought will be enormous.
Oasis Capital: Research sometimes is about breaking through window paper. In robotics, assuming we want to break through from the present to NLP's "phenomenon-level" moment, does academia still need to承担 more of the role, or have we already reached the engineering node?
Professor Yin: This is a great question. Currently, we still can't do without either academia or industry. How to understand this? Look at Tesla's autonomous driving FSD. Tesla FSD's Transformer came from pure Google researchers. Whether OpenAI, Google Research, or Meta, although their work is also industry-oriented, they have sufficient resources in an industrial environment to support research work.
What's interesting about Tesla is that although it uses very old Transformers or BEV Transformers, it can构成 the data闭环 of the autonomous vehicle industry. Everyone talks about data闭环; pushing pure algorithms to the extreme also has boundaries. FSD v12 is a典型案例, showing everyone how system performance can keep improving. It has a general architecture that can run, massive high-quality high-performance data, and this data can repeatedly optimize the system, giving the model sufficient generalization ability.
FSD's strength lies, on one hand, in having a good enough architecture, and on the other, in being able to form an effective high-efficiency data闭环 under a good enough architecture, getting the model flywheel spinning to reach ideal results. ChatGPT is the same — architecture flows through, massive high-quality data emerges, completing the reinforcement learning闭环, and the flywheel spins.
From these two examples, we can see that effective emergence of ultra-large models on the real world requires architectures from academia (academic thinking is more open and active, forming more effective reinforcement learning mechanisms or architectures, so architectures generally come from academia), while simultaneously requiring efficient data integration capabilities from industry. Tesla is undoubtedly typical — millions of vehicles worldwide collecting data simultaneously, encountering many corner cases (long-tail scenarios), which is simply impossible in academia.
Another example: around 2018-2019, OpenAI's CTO wanted to build a robotics large model, using a robotic hand to solve a Rubik's cube, but generalization ability was extremely poor and the project was eventually paused. Afterwards, the CTO stated in an interview that the biggest短板 of robotics models is how to obtain high-quality effective data. This reflects that whether academia or top institutions, they can indeed推出 architectures, but whether the model can spin and continuously optimize depends on continuous quality data.
Whether NLP or autonomous vehicles, we now have clear templates to push forward. But in robotics, whether PaLM-E or VoxPoser, data dimensions are too complex — auditory, sound, tactile, spatial, plus perception. How to couple such multi-dimensional information into a multi-dimensional database for training has not yet achieved breakthrough.
Oasis Capital: Speaking of Tesla, do you think Elon Musk's bet on pure vision back then was first principles? Or did he see the possibility of ultra-large models?
Professor Yin: That's really hard to say (laughs). Musk's thinking mode — on one hand it may indeed be first principles, but on the other hand there's also gambling involved. Musk's conclusion that lidar is useless and pure vision is feasible was based on "this is how humans do it," which fits first principles, but there's undoubtedly also an element of gambling. Opponents can easily say: humans do indeed do it this way, but airplanes don't have bird structures — human inventions enable airplanes to achieve bird capabilities. From this standpoint, first principles doesn't hold up either.
But first principles must also be considered from another底层. Lidar costs are indeed too high; if vision can achieve it, cars can be cheaper, sales volume goes up, and data volume can rise. According to large model flywheel mechanics, Tesla's models will inevitably train better and better. Going back 10 years, it was hard to evaluate whether he was gambling or truly saw something. In 2013, deep learning was just emerging, and it was hard to determine whether avoiding high-precision lidar was feasible. Looking at the present, what everyone sees is that whether from production or car-selling perspectives, regardless of whether general artificial intelligence or general vehicles can be achieved, as long as costs can be reduced, volume rises, and data comes in, everything else becomes manageable.
Oasis Capital: Do you think future底层 control for robots can be standardized? Including the operating system you work on?
Professor Yin: This can also be understood from history. When airplanes were first born, there were multiple modalities — glider forms, bicycle-pedaling with wings attached, a dazzling array. With the birth of the earliest batch of airplanes, due to aerodynamics, high-efficiency low-cost practical demands, they eventually converged into several conventional modalities. Robotics is the same — different control methods correspond to different robot bodies, humanoid, wheeled-legged, legged, bipedal, quadruped... Ultimately, based on customer needs, they will converge into one or two conventional modalities.
From the customer needs perspective, why is everyone working on humanoids now? Because under conventional scenarios, converging to humanoid is a relatively stable state. Will all robots converge to humanoid? Not necessarily. In certain environments, wheeled-legged robots have advantages in power consumption and movement efficiency, and this type of robot载体 may also take shape.
When all robots converge to several conventional modalities, this will only affect the底层 control layer; it won't affect anything above the control layer. Conventional control strategies will also eventually converge into one or two. Taking control methods as an example, early on there was PID control, MPC control, H-infinity control... In 2019, ETH's quadruped robot developed a set of end-to-end training methods based on reinforcement learning, mapping from simulation to reality. The biggest advantage of this method is using limited real-world data plus massive virtual world data, through a set of reinforcement learning methods, to train the quadruped robot exceptionally well. Their latest videos also show stunning results, and this strategy was验证 successful back then.
Recently I've heard that some wheeled robots or wheeled-family robots are also using similar strategies, through end-to-end methods, to train this type of robot to a stable state.
This example shows that control methods, regardless of what form the底层 robot takes, will eventually converge to one or two control strategies. From the operating system perspective, after底层 fine-tuning, it can also适配 different modalities of robotic systems. This may be a future trend.
Oasis Capital: Following this logic, a sufficiently general-purpose robot, in 10 or 15 years, could theoretically really do everything?
Professor Yin: It probably won't take that long. Around 2030, conventional general-purpose robots should basically be able to enter people's lives.
Oasis Capital: Everyone is exploring various paths to embodied intelligence, with many papers recently. Benchmarked against AI, has a通道 already emerged for embodied intelligence in robotics? Or is there currently no consensus?
Professor Yin: Whether domestically or in the United States, the entire industry is gradually forming consensus around humanoid robots. Take Tesla — besides making humanoid robots, they also do SpaceX, which just had its second launch on August 31st. For them, their internal established timeline is around 2028-2030 to send both humans and humanoid robots to Mars. To achieve this goal, humanoid robots must possess basic capabilities for working in space and Mars environments.
In 2021 Tesla announced they would make humanoid robots; last year they had初级 modalities, and this March at AI Day they demonstrated basic gait control — the basic form has been polished. From inside out, control machines, motors, all with explained design rationale. In 2008 Tesla wanted to build electric vehicles; back then few people in the world responded, and now the entire industry has been changed by him alone, and China's entire electric vehicle industry has been changed too.
Building cars is harder than building robots — robots are nothing but motors and actuators. When all these things can be quantified, Musk breaks down all complex matters, filtering out production costs layer by layer, achieving low-cost落地. Musk believes that by 2028, humanoid robot production efficiency will reach about 1/3 of human efficiency, with robots working 24 hours unlike humans' 8 hours. Plus data flywheel, standardized modules and operation modes, after accumulating massive quality data in factory environments or specific environments, robotic FSD won't be a problem.
Robotic data is higher-dimensional and more complex. Between 2023 and 2030, quality companies both domestically and abroad are trying various ways to optimize humanoid robot or other form robot operational modality data into a database, using similar large model systems for continuous optimization training, the same method as FSD.
Is this the right通道? Synthesizing Tesla's history of rising up, from declaring they would achieve ultimate autonomous driving to the intermediate process now realized, plus OpenAI's progress, we can sense that whether this works or not depends on several key steps: first, whether robot body, data, and structure can be standardized? Second, whether large models can advance further? If all these are achieved, this time point could be realized.
Oasis Capital: You mentioned a key point — "data." The entire industry brings up this issue. What are your thoughts on data? What explorations are people doing?
Professor Yin: Still taking Tesla FSD as an example, everyone can have an intuitive feeling. Why is its FSD good? Simple — it has massive numbers of cars,梳理 massive corner case data. Why can't other companies match it? Because general automakers' data collection is too subjective and not passive. This leads to no matter how elegant and graceful the method design, the data carries bias. People objectively seek good scenarios to drive in, ignoring corner cases. Why can Tesla do it? Because they have a standard data synchronization mechanism — no matter where in the world they're driving, more of it is accident data, i.e., abnormal data. Tesla continuously runs and updates models in the cloud — this is FSD's true core of growth.
When robots can spontaneously collect data without bias, then this can be accomplished. Of course cars and humanoid robots are different — cars have strong universality in standard highway environments. Robots, especially humanoid robots, need to adapt to complex scenario modalities. Only when this mechanism is established, achieving spontaneous robotic data collection, recording corner cases, continuously maintaining cloud databases, and then optimizing robot behavior does this become possible.
Oasis Capital: People mention simulators, or simulation-to-reality platforms whose physical performance currently seems insufficient. How well do NVIDIA and other open-source tools match with robots? In this direction, do startups have opportunities?
Professor Yin: I think it's relatively difficult. A startup doing autonomous driving data might make sense, but actually not many companies have truly跑出来 — most still use real data. The biggest problem with doing it through simulators is that data modalities are too complex, and the machine's own control variables are complex, making collecting high-quality data in the real world basically a伪命题.
High-quality platforms already exist on the market — Unreal, Unity both do quite well. If a startup wants to build its own simulation platform, the time cycle and cost relative to major manufacturers who have been in the industry for over a decade will present some difficulties.
From another perspective, there are also companies like ETH's quadruped robot, which can run so robustly because it's highly bound to NVIDIA. NVIDIA provides platforms that approximate the real world as closely as possible, putting the quadruped robot in for verification, only needing to collect limited real-world data to generalize massive virtual-to-real data. From this思路, cooperating with major manufacturers is feasible.
Of course, it doesn't rule out individual very strong teams that can build powerful physics engines themselves — this has happened in history. Physics engines are still critical; rendering effects alone aren't enough. If a team achieves breakthrough in physics engines, they might carve out a place. But currently, whether open-source physics engines or rendering projects, it's hard to see very valuable content.
Oasis Capital: Cooperating with major manufacturers, besides providing open functions for everyone, what specifically will NVIDIA add?
Professor Yin: This is bidirectional. NVIDIA Omniverse also does experiments, just like when NVIDIA did CUDA. In CUDA's early days, there was simply demand from everyone, letting NVIDIA judge on its own how this thing should develop. It optimized tools based on feedback from the academic community or industry. Similarly, in the robotics industry when Jensen Huang said to go all-in on embodied intelligence, he was planning another tool similar to CUDA. As for what demands the industry actually has, they advance based on everyone's feedback. On the other hand, the industry will use integrated toolkits to further rapidly optimize systems, mutually弥补 each other's shortcomings.
Oasis Capital: Under robotics' current development, what are industry's actual demands for robots?
Professor Yin: The industry I接触 is AMR (Autonomous Mobile Robots) or mining area robots. Taking mining areas as an example, from the customer's perspective, mining areas actually need a complete robot system to fully replace humans: automatically entering underground mining areas, environmental modeling, quality analysis, generating data needed by mining area users, arranging maintenance. They need integrated solutions, not a single device. Similarly, taking logistics as an example, logistics robots only place goods on delivery lockers — there's still a gap to getting them into people's hands, missing the stairs and entering the campus. What people need is completely eliminating humans and delivering things to users' hands.
Oasis Capital: You previously mentioned that around 2030, conventional general-purpose robots could enter people's lives. Which robots does this refer to?
Professor Yin: The factory environment where Musk's humanoid robots operate is already a very complex scenario. If humanoid robots can handle factory environments with ease, then in principle from a technical perspective, this type of robot can directly enter society. Including public safety, such as street patrol robots,扫地 robots; home service robots, caregiving robots. By 2030, it's expected that this type of robot will replace all currently known扫地 cleaning, security, logistics, and patrol robots, leaving only one modality of robot.
Of course the hardest part of the process is that even putting robots in pure factory environments, after 5 to 7 years of iteration, how to give them adaptation capabilities in various environments? For example, can they easily navigate through dynamic crowds? In noisy factory environments, can they focus on completing work while avoiding collisions with other workers? Can they receive multimodal information, such as factory environment alarms, and take optimal actions? If this series of problems can all be solved in factory environments, then this type of robot can basically enter thousands of households.
Oasis Capital: What was your original intention for choosing to work on robots?
Professor Yin: Human physical capabilities are limited. Humans want to go to outer space, underground, underwater. If a robot with labor capabilities, overcoming environments, and working in diverse scenarios can be born to replace humans, then human activity range, activity capabilities, and human productivity will produce completely different conclusions. Like most people working on technology, I hope to improve productivity.
Oasis Capital: To what extent can your current research work solve real problems?
Professor Yin: For example, high-precision localization, achieving automatic excavation, automatic exploration, automatic optimization, ensuring robots can work continuously for years without stopping — this is what we can currently do. The general robotic model we've built also further optimizes on the basis of inheriting the above work. This type of robot, without any human intervention, can continue to survive in natural environments and maintain exploratory capabilities.
Oasis Capital: What was your original intention and vision for founding MetaSLAM?
Professor Yin: To bring together the best robotic top-level technologies, unite world-leading robotic technologies, and build a general robotic system — this is the purpose of founding this organization. Technologies can be mutually通用 and shared within the organization, hoping to gather everyone together to create an open-source infrastructure operating system.
Celebrating Vitality
What do you think is technological vitality?
The only unchanging law in this world is "continuous change" — only through constant internal and external evolution can we adapt to the "new vitality" brought by technology. — Professor Yin Peng
Assistant Professor, City University of Hong Kong


Oasis Capital is a new-generation venture capital firm in China, dedicated to discovering the most vital entrepreneurs of the next decade and growing alongside them to create long-term value. "Celebrating Vitality" is Oasis's vision and mission. This vitality is both the direction of era-defining structural transformation and the resilient, evolving power of entrepreneurs.
Oasis Capital focuses on early and growth-stage investments, with individual investments ranging from $3 million to $30 million USD, concentrating on robotics, artificial intelligence, technology services, and other fields, empowering China's technology-driven new service upgrade.



