The Moon and Sixpence of Embodied Intelligence

Counselor Vitality

"You can count the seeds in an apple, but you can never count the apples in a seed."

Technology works the same way.

Too often, we see technology as a result — an apple. But technology is actually a seed, a beginning. Oasis Capital believes in the societal transformation that AI will bring. We believe even more deeply that AI is just one cross-section of humanity's entry into the intelligent age, and embodied intelligence is another.

It is this conviction that led Oasis to accelerate infrastructure investments in embodied intelligence during the market downturn of the past two years. So starting this year, alongside our global AI expert dialogues, we look forward to sharing more of our understanding and observations on embodied intelligence with you. We hope for rich exchanges with all of you — so we can grow more apples together. Part I

The AI Leap: Robots Enter a New Development Era

Building general-purpose intelligent robots that operate from a first-person perspective, achieving closed-loop autonomous perception - planning and decision-making - autonomous execution, and adapting to diverse scenarios has long been the ultimate goal in robotics and AI. Traditional robotic systems have consistently fallen short of these capabilities. Since GPT-3's debut, AI's explosive growth has yielded remarkable achievements. In recent years, the emergence of various LLMs and LVMs has brought new hope to robotics: the concept of embodied intelligence was born — intelligent systems capable of understanding, reasoning, and interacting with the physical world. For AI, embodied intelligence represents the critical vehicle and entry point for general AI to interact with the physical world; for robotics, embodied intelligence will lead the field into a new era of generalization.

From Past to Present

Looking back at robotics development, the essence has been a progression from specialized machines toward greater generality, and from passive programmed control to active decision-making. In the first two generations, robots functioned more like specialized automation and intelligent equipment — solving targeted tasks through specific mechanical structures for particular scenarios. They suited relatively simple, fixed, structured environments with very limited generalization and transfer capabilities. Moreover, every capability required meticulous engineering programming. Because robots lacked deep understanding of task objectives, engineers had to perform extensive task decomposition and programming, heavily relying on hand-written code for robot control. Each step needed precise human planning, and any change to the task object or environment demanded complete reprogramming and redeployment — time-consuming and labor-intensive.

Robot 1.0 Era Robots primarily served automation goals, widely deployed in industrial settings for precise, repetitive assembly and handling tasks. Forms included industrial arms, small six-axis robots, SCARA arms, and others executing along preset trajectories and speeds. Mobile platforms were represented by traditional industrial AGVs navigating via fixed beacons. Robots at this stage possessed no perception or decision-making capabilities, relying entirely on pre-programmed control instructions for fixed production line tasks. The "Big Four" robot families dominated this core market. New opportunities mainly came from import substitution drivers and downstream demand from terminal industries like automotive. Domestic brands including Efort, ESTUN, and New Era emerged on stage, pushing the industry's overall automation capabilities forward.

Robot 2.0 Era

Robots began acquiring preliminary perception and planning capabilities, evolving toward greater intelligence. Breakthroughs in discrete technologies and new product category innovations drove several waves of entrepreneurship and new opportunities in the Robot 2.0 era:

  • SLAM Technology Drives New Mobile Robot Form Factors

    As SLAM spatial perception and localization technology matured, its integration with robots endowed them with autonomous mobility, opening multiple new product category tracks. Roborock pioneered applying SLAM to household robot vacuums, transforming the market from random bump navigation to intelligent path planning — directly disrupting the robot vacuum landscape. In industrial mobile robotics, where traditional systems relied on magnetic strips or QR codes for navigation, the combination of SLAM and LiDAR opened new possibilities for next-generation AGVs (Automated Guided Vehicles) and AMRs (Autonomous Mobile Robots) with autonomous mobility. Driven by downstream demand from e-commerce warehousing, dozens of startups emerged, giving rise to companies like Geekplus, Quicktron, and Hai Robotics. In service robotics, autonomous mobility built on SLAM enabled delivery, reception, serving, cleaning, and other mobile service robot products. Companies like Gaussian Robotics, Keenon, Pudu, and Yunji each carved out markets in service robotics.

  • Collaborative Robot Arms as a New Product Category

    In 2012, Universal Robots, the world's first lightweight collaborative robot arm manufacturer, entered the China market, introducing collaborative robots to the country. Collaborative arms, with their lightweight, compact form factors, relatively lower precision requirements, and simple interaction interfaces, could operate in human-robot collaborative scenarios — becoming an emerging product category in the robotic arm space. Since collaborative arms started later, the gap with other high-precision industrial arms was smaller, seen as an opportunity for domestic manufacturers to leapfrog. Shortly after UR's China entry, and driven by Industry 4.0 and Made in China 2025 initiatives, a collaborative robot entrepreneurship boom began around 2015. Companies like AUBO, Jaka, DOBOT, and珞石 emerged in response.

  • AI + 3D Vision Technology Drive

    In the robotic arm domain, AI and 3D technology gave robots visual perception and intelligence — arms gained "eyes." Through AI combined with 3D vision, robots could automatically recognize and localize objects, achieve optimal path planning, reducing dependence on manual teaching and deployment. Hand-eye coordination flexibility expanded industrial robot application scope, successfully solving non-standard automation challenges like loading/unloading, depalletizing, random bin picking, and welding — bringing new opportunities and market space.

Robot 3.0: General Intelligence Stage

The generalization capabilities demonstrated by large models have brought entirely new possibilities for general intelligent robots. Unlike the discrete, point-solution drivers of the previous stage, we believe this AI-robotics integration will comprehensively reconstruct the robot's overall system capabilities across perception - decision-making - control, expanding robot capability boundaries more broadly, creating more extensive market opportunities, and impacting the robotics industry more comprehensively and profoundly. Robotics enters an entirely new development paradigm.

Part II

Deconstructing Robot Foundation Models

Research at the frontier of combining AI large models with robotics has become a major hotspot in academia and industry, with results continuously emerging (see figure below). So how will robot foundation models reshape the robotics development paradigm?

A Robotics Systems Perspective

The core system components for robot operation include: Perception, Decision Making and Planning, and Control & Act Generation. In the embodied intelligence framework, these correspond to high-level brain decision-making and low-level cerebellar control and execution systems. The "brain" handles task understanding, decomposes and plans tasks combined with perception information, and formulates execution strategies. The "cerebellum" handles core motion control, implementing robot action execution and feedback under the brain's strategic direction.

We categorize large models integrated with robotics into two types: Foundation Models for Robotics and Robotics Foundation Models. The former comprises various LLMs, VLMs, and VFMs that can integrate with robotics but are not limited to robotic applications. The latter are foundation models trained on robotics data that can generate down to the cerebellar control level — that is, robot embodied foundation models.

Multimodal Foundation Models and the Robot Brain

For the first category — language LLMs and visual multimodal VLMs — these act at the robot brain level. Such large models provide robots with powerful general understanding capabilities and strong interactivity, while incorporating human society's knowledge and common-sense systems, enabling high-level task abstraction and understanding-based planning — substantially enhancing robot brain capabilities.

  • Perception Level: Large Models Enable Multimodal Perception

    Traditional robot perception systems often relied on single data sources or sensors, and were constrained by conventional AI capabilities — requiring massive data annotation and supervised learning for each different object, with weak generalization. Multimodal foundation models can fuse heterogeneous multimodal data from different sensors, learning and understanding text, images, video, audio, and other modalities for higher precision. Their strong generalization capabilities allow robots to achieve more general perceptual recognition through fine-tuning on small amounts of new data. A representative example is the PaLM-E model, integrating the 540B-parameter PaLM with the 22B-parameter visual ViT model, extending large model capabilities to computer vision and providing robots with general visual perception capabilities.

    It is foreseeable that robot perception systems will integrate more physical world dimensional information in the future, such as force control, touch, olfaction, and physical laws. Multimodal models will evolve more deeply into multidimensional world models, endowing robots with richer and more precise multidimensional perception capabilities than humans possess.

  • Planning Level: Large Models Replace Engineers for High-Level Planning

Before large models, application engineers spent most of their time understanding tasks, breaking them down into appropriate actions, and writing, tuning, and deploying robot applications using robot programming languages. Traditional robot control relied on precise modeling — but such models were typically built for specific environments. Any change to the environment required rebuilding the model, leaving limited transferability. The high-level abstraction capabilities of large models mean that the work engineers once spent enormous amounts of time on — task definition, decomposition, and programming — can now be directly handled by these models, enabling robots to truly achieve autonomous task planning.

Multimodal large models can be integrated with robotic applications to enhance the robot's "brain," but simply prompting off-the-shelf LLMs/VLMs to directly control robots remains challenging. These models are not well-suited for low-level precise control and still require supplementation from a "cerebellum" layer of control capability. Therefore, how to better align the brain's decision-making and planning with the cerebellum's execution — achieving a closed task loop from large model to robot — is critical and a problem that research teams are currently working to solve.

Embodied Large Models and the Robot Cerebellum

Leveraging multimodal large models can provide high-level planning for the robot's brain layer, but they cannot directly achieve low-level control and action generation. Relying solely on brain models therefore cannot complete the final closed loop of robotic execution. By contrast, the cerebellum layer serves as the underlying control system for controlling and executing actions, translating high-level decisions into concrete action execution and bearing responsibility for final outcomes. Robotics Foundation Models — embodied models — extend into cerebellum-level control, combining robot data training to achieve generative capabilities at the robot's action execution end. Representative examples include Google's RT series, which uses transformer models combined with real-world collected robot data to achieve output from raw inputs such as images and speech to end-effector actions, demonstrating certain generalization capabilities. However, limited by the scarcity of robot data, the current data and parameter scale of embodied models still falls significantly short of building true embodied large models. Taking manipulation as an example, better cerebellum control requires building a rich library of primitive-level actions at the foundation, such as grasping, wiping, folding, and placing. Unlike brain-layer model training, which can be decoupled from specific hardware form factors, the cerebellum layer requires strongly coupled training between algorithms and hardware, with substantial action data. Therefore, in cerebellum embodied models, skill learning becomes the foundational task for achieving embodied intelligence; the success rate training and generalization of skill sets are key issues, and learning some complex basic skills remains a major difficulty.

- Robot Low-Level Control: From Model-Based to Learning-Based

Traditional robots typically employ classical control strategies, achieving basic motion control through direct drive or motor control. This approach primarily involves dynamics modeling and constraint design, with PID controllers used for regulation at the lowest level. The integration of AI with control — through imitation learning and reinforcement learning, among other methods — is increasingly adopted, gradually shifting robotic control toward autonomy and more advanced adaptive and intelligent control. Taking the development history of legged robot control as an example, it has roughly gone through three stages:

Stage One: Rule-based simplified models dominated by LIMP (Linear Inverted Pendulum Model) + ZMP (Zero Moment Point). This approach simplified position control but suffered from poor gait stability.

Stage Two: Introduction of dynamics models, MPC (Model Predictive Control) + WBC (Whole Body Control), enabling dynamic modeling and control that supported more diverse movement capabilities. However, the algorithms were sensitive to hardware variations, and complex, precise modeling remained difficult.

Stage Three: In recent years, AI-integrated control algorithms incorporating deep learning and reinforcement learning. AI control offers strong environmental adaptability, breaking through limitations of traditional control methods and raising the ceiling of robotic movement capabilities.

Model-Based Control offers advantages in stability, interpretability, and real-time performance, ensuring the floor of robot control capability in simple control tasks. Learning-Based Control, through data-driven methods, shows potential in solving complex and high-dimensional motion control problems. However, due to its black-box learning characteristics, interpretability, reliability, and safety still need improvement, and the training process requires substantial data and computational resources. In the legged control domain, classical model control and AI control are fused and applied for different task scenarios; in embodied manipulation and general control domains, current research priorities trend toward learning-based control strategies to achieve more intelligent and generalized robot manipulation.

Part Ⅲ Exploring the Implementation Path of Embodied Intelligence

End-to-End Models vs. Hierarchical Decision Models

On the path toward embodied large models, two mainstream approaches currently exist: end-to-end embodied models (represented by Google RT-2) and hierarchical decision models (represented by Figure 01).

End-to-end models complete the entire process from task goal input to direct control signal output through a single neural network. This approach is represented by Google RT-2, whose goal is to train an end-to-end model that learns from robot observations to actions, leveraging large-scale pretrained VLM models to directly generate low-level robot motion commands. RT-2 first pretrains VLMs on large-scale internet data, then fine-tunes on robot tasks, combining robot action data to introduce the Vision-Language-Action (VLA) model, enabling end-to-end output from images to control commands. The RT-2 model also demonstrated better generalization capabilities. End-to-end solutions can universally and automatically adapt to various environmental changes, appearing to be the most direct and ideal implementation direction. But in practical deployment, they face numerous difficulties: end-to-end approaches require massive data for training and consume enormous computational resources. And the larger the data scale, combined with high-frequency large model calls, the more robot decision speed degrades, affecting real-time performance. Of course, with sufficient computing resources and data, the brute-force approach enabled by large-scale data and computing power could genuinely produce stunning results.

The hierarchical decision model approach decomposes tasks or modules into different levels, training multiple neural networks and combining them in a pipeline fashion. For example, Figure AI's Figure 01 links to an OpenAI large model at the top layer (possibly GPT-4V), providing visual reasoning and language understanding; the middle layer consists of Neural Network Policies (NNP), where neural networks serve as the cerebellum for low-level control, generating a series of fast, low-level, dexterous robot actions that directly map pixel information to action commands at extremely high frequency. Figure 01's policy network can output action commands at a control frequency of up to 200Hz — an impressive feat that end-to-end approaches would struggle to match in responsiveness. The bottom layer is Whole Body Control, which receives NNP action commands and executes the lowest-level control. This hierarchical architecture is relatively controllable, simpler to implement, and can be broken down into different small models at the bottom layer with good interpretability, making it a more suitable choice when early data volume is insufficient. But distributed decision-making requires more time to solve alignment and consistency issues between different steps.

Imitation Learning and Reinforcement Learning

Training methods on the embodied intelligence cerebellum side mainly concentrate on two categories: imitation learning and reinforcement learning.

- Imitation Learning

By observing an expert (human or another machine learning model), this approach tends to rapidly learn from skills demonstrated by excellent performers. This learning method is essentially a direct mapping of demonstrations, so the learned policies and behaviors are typically bounded by the demonstration data and cannot exceed the capability boundaries shown by that data. The advantage of imitation learning is that it is relatively direct and simple, enabling rapid acquisition of knowledge and skills from expert demonstrations.

Representative examples include Stanford Aloha, which provides a complete teleoperation system for imitation learning that can very quickly learn seemingly complex and long-horizon task combinations. However, imitation learning's limitations lie in its upper bound being the experience and strategy of the imitated subject; its transfer and generalization capabilities are relatively weak, and it may not effectively adapt to changing environments or task demands. It also requires large amounts of high-quality demonstration data to ensure learning effectiveness and model performance stability.

- Reinforcement Learning

- Reinforcement Learning

Reinforcement learning (RL) designs reward mechanisms that enable robots to learn how to maximize cumulative rewards through interaction with their environment for specific tasks. Unlike traditional supervised learning, RL doesn't rely on pre-labeled data; instead, it discovers optimal behavioral strategies through trial and error. If imitation learning is like copying the top student, RL is about becoming the top student yourself.

RL excels at decision-making in complex environments, allowing robots to explore autonomously with strong generalization capabilities across diverse environments and tasks. However, designing appropriate reward functions is critical to RL's effectiveness, and reward engineering for certain long-horizon tasks remains challenging. RL also demands enormous sample sizes, and training on physical hardware carries high damage costs. Consequently, most current RL research relies on physics-based simulators for training, making the sim-to-real transfer problem a key hurdle to overcome for practical deployment.

Imitation learning and reinforcement learning are not mutually exclusive approaches; they can be effectively combined based on specific scenarios and requirements. For instance, imitation learning can serve as a warm-start for RL—by mimicking existing expert policies, robots can begin task execution more quickly, then gradually transition to the RL phase without exploring the environment from scratch. For structurally simple, easily modeled tasks, RL in simulation environments enables rapid training. For long-horizon tasks or complex scenarios beyond simulation capabilities, imitation learning offers a more direct and effective path. The two approaches complement each other, leveraging their respective strengths to improve both training efficiency and performance in embodied intelligence systems.

Real-World Data vs. Simulation Data

Scaled data is foundational to current embodied intelligence models, with different training approaches emphasizing distinct data collection pathways. Real-world data collection and simulated environment data constitute the two major categories of embodied training.

- Real-World Data

Data collected directly through robot hardware in physical environments represents the highest quality and most immediately usable form—but also the costliest. It demands substantial human labor and hardware investment, and physical training carries significant damage risk, making large-scale collection in real environments extremely difficult. Google, leveraging its financial and research capabilities, spent 17 months collecting 130,000 real-world robot trajectories across 13 robots, laying the data foundation for RT-1 and RT-2.

Teleoperation: Direct human control of robots for data collection yields high-quality, immediately usable data. The drawback is the required investment in labor and hardware for real-world collection. The better implementation of teleoperation involves building systems that accommodate diverse robot platforms, integrate multiple sensors, and operate at lower cost with greater safety.

Real-hardware learning is direct and high-quality but difficult to scale. This has led to indirect methods of human data collection.

Motion Capture Data Collection & Video Learning: These approaches don't collect data directly through robot hardware, but instead learn from more readily available human behavioral data—through wearable devices or human behavior videos. The advantage is lower cost; video learning can leverage abundant free internet data at scale. However, the core challenge is mapping data from humans to robots. Such data is difficult to process, noisy, and of limited effectiveness. This passive data can be used for task pre-training; better solutions to the mapping problem could substantially alleviate data bottlenecks.

Additionally, Apple's Vision Pro launch offers new possibilities for device-based collection. If Vision Pro serves as a data collection entry point, its growing penetration could make it a more ubiquitous human data collection platform.

- Simulation Data

Data collection through simulation simulators represents an alternative path to real-world physical data. Simulation data is low-cost, obtainable at scale, and enables 24/7 continuous training. Complex control problems like locomotion have achieved notable training results recently through simulation data combined with RL—the impressive quadruped robot demos we see are almost all trained in simulation. Common simulators in robot RL training include Isaac Gym, Isaac Sim, ManiSkill, and MuJoCo.

However, simulation data quality depends heavily on simulator capability. Modeling discrepancies in environment, perception, robot hardware, or control within simulators can all produce sim-to-real gaps, making this reality gap a persistent challenge requiring continuous attack. Additionally, simulating soft objects like shoelaces and clothing remains an unsolved difficult domain, potentially limiting simulation data's applicability to soft-object manipulation. Long-term improvements in more accurate and efficient simulation performance will represent a major breakthrough for the robot data challenge.

Whether based on simulation or real-world collection, whether using RL or imitation learning pathways, different data collection and model training methods form a diverse skill tree for embodied intelligence training. Currently, various teams leverage their respective strengths, exploring directions from simulator performance optimization and teleoperation hardware-software improvements, to reward mechanism refinement and innovative data collection methods, all working toward building models with stronger generalization. Throughout this process, how to deeply optimize skill tools, how to select matched training pathways based on task understanding, and how to allocate time and resources across different branches of the skill tree will all be core factors affecting embodied model training outcomes.

Part IV

Mining the Embodied Intelligence "Sixpence"

Three Core Requirements for General-Purpose Embodied Intelligence

New embodied intelligence achievements are emerging globally, bringing considerable confidence and anticipation for general-purpose embodied intelligence. Yet while gazing at the "moon," we also recognize present challenges. Three core elements underpin general-purpose embodied intelligence: scaling laws for robot algorithms, a general-purpose robot hardware platform, and large-scale data flywheels.

- Data Flywheel

High-quality training data for robots remains severely scarce. Both simulation and real-world data face unresolved challenges; meeting the data volume requirements of large models still demands prolonged collection and accumulation. The data side represents both the greatest challenge and core moat for achieving embodied foundation models and robot generalization. Building an efficient, high-quality, low-cost data training-feedback mechanism is therefore critical.

- Scaling Law

Where is the scaling law for robots? Brain-side perception and planning have already gained substantial generalization improvements through multimodal foundation models. The embodied model scaling law involving robot motion control is what the field is currently exploring. Achieving scaling laws requires both better "recipes" and sufficient "ingredients." On the recipe side, our earlier analysis covered different training pathways and perspectives for embodied intelligence. As embodied research deepens, algorithm models may achieve further breakthroughs in generalization and universality, but the pace will remain constrained by dataset scale. Progress toward scaling laws will more likely be modular and phased.

- General-Purpose Hardware Platform

Hardware is the ultimate execution vehicle for robots. How do we define a general-purpose hardware platform that, within acceptable cost constraints, can handle diverse task requirements, execute reliably and stably, facilitate efficient data collection and feedback, and ensure adequate safety? Hardware design iteration and performance optimization will be a continuous grinding process.

Understanding Embodied Intelligence from a Robotics Industry Perspective

Embodied intelligence remains fundamentally within the robotics industry chain, following robotics industry development patterns. Looking back at robot development history, compared to the AI industry, the robotics value chain is more complex and lengthy. On one hand, algorithms and hardware are tightly coupled; on the other, the journey from frontier technology breakthrough to product demo formation, through PMF stage into small-batch production, to large-scale commercial application requires a longer landing cycle. From an exciting demo to actual product mass production and commercialization involves numerous engineering challenges around stability and robustness—demanding sufficient patience and resilience.

Robot core value dimensions manifest in three aspects:

- Extending human capability boundaries: Executing dangerous, extreme, or humanly difficult tasks

- Labor substitution: Replacing humans in repetitive, monotonous, or high-intensity work

- Providing emotional interaction and companionship

In the first two value dimensions, robots serve as new productivity tools, and application decisions must still follow ROI calculations weighing human labor costs against robot replacement costs. Therefore, regardless of the technical halo surrounding embodied intelligence products, they must still face the core constraints and challenges of robot industry scaling: demand validation, cost control, customer success, hardware supply chain management, and delivery and fulfillment efficiency. Respect the industry development chain, maintain patience and resolve, and keep your feet firmly on the ground even as you gaze up at the moonlight.

The embodied intelligence field is currently experiencing rapid technological innovation and breakthroughs — advances that will greatly expand the potential value and application boundaries of robots across scenarios, creating comprehensive market opportunities. But technology breakthroughs alone are insufficient to support the industrialization of embodied intelligence. Embodied intelligence breakthroughs must be combined with sound functional or product definition and real scenario demands to truly achieve commercial viability. When determining application scenarios, technology is only one part of defining capability boundaries. How to combine those technical boundaries with scenario-specific product definition requires careful deliberation and judgment of genuine needs — digging deep for the "sixpence" in embodied applications.