suho x Gaoyang: Embodied Intelligence Is the Only Answer for AI Entering the Physical World | Agentic Era

When Intelligence Enters the Physical World

The Agentic Era was one of the central themes at Oasis Capital's AGM this year.

As AGI accelerates toward reality, Embodied AI is transforming from a technical subfield into a critical gateway for understanding general intelligence. As model capabilities advance, the question of how to bring intelligence into the physical world — to learn and act in real environments — has become one of the most widely shared imperatives of the Agentic Era across academia and industry.

For this year's AGM, we invited two founders from the Oasis portfolio who are also among the most representative scientists in this field:

suho, tenured professor at UCSD and co-founder of Hillbot, one of the earliest scholars to propose the term "Embodied AI" and a driving force behind its emergence as an independent discipline.

Yang Gao, assistant professor at Tsinghua University and co-founder of Spirit AI, who conducted extensive research in computer vision at UC Berkeley and brings rich hands-on experience in embodied foundation model training and real-world robotics experimentation.

Both speakers have long operated at the frontier of embodied intelligence research and commercialization, with deep insights into technical evolution, model architectures, China-US competitive dynamics, and inflection points. This conversation centers on several core questions:

  • Is embodiment a necessary path to AGI?
  • Where do the real technical bottlenecks lie?
  • What structural advantages does China hold on this path?
  • When will the next perceptible leap arrive?

The following is an edited transcript of approximately 8,000 words, with an estimated reading time of 20 minutes.

Enjoy

Oasis Capital: I'd like to start by asking both of you to share how you began your entrepreneurial journeys.

Professor Yang Gao: Let me begin with some reflections from my time at Berkeley.

I started my PhD in computer vision, then shifted to robot learning around two years in. Around 2016, those of us at Berkeley had this vague intuition that robotics represented the ultimate, shining jewel of intelligence — but we didn't really know how to approach it. We just sensed it was the crucial last mile for solving intelligence.

OpenAI had already been founded, with some faculty members like Pieter Abbeel working there part-time. OpenAI was already preaching "scale compute for miracles," which many Berkeley professors dismissed as unreliable. This went on for about three years, until GPT-3.5 emerged. I think that's when some people began to shift their thinking — they saw something different.

By then it was around 2021, and I had already returned to China.

At first, I didn't think much of it. Everyone took for granted that "more data, more human-like model outputs" was unremarkable. But when OpenAI introduced post-training and RLHF (Reinforcement Learning from Human Feedback), that's when I believe many senior professors at Berkeley, myself included, experienced something like a collapse of faith in how AI research should be done. We realized OpenAI had been right all along. That was probably the most transformative intellectual moment for me — realizing this actually worked, that such a simple approach of scaling data could solve it.

Why am I building an embodied intelligence startup now? It was precisely at that moment I realized the underlying logic is identical for embodied intelligence. The only difference is that instead of just producing speech, we're dealing with actions — systematic responses to diverse input modalities including vision, language, touch, sound. Fundamentally, there's no distinction.

Working from first principles, regardless of timeline, embodied intelligence will experience massive leaps from the exact same technological trends and extrapolations within a few years.

So at that point, I knew embodied intelligence was coming, and I wanted to be the one to do it.

Oasis Capital: Thank you, Professor Gao. And Professor suho?

Professor suho: I entered AI even earlier, around 2002, long before the deep learning era.

My advisor at Beihang University was a mathematical logician, and everyone was researching automated theorem proving. We soon discovered the boundaries of rules — expressing common sense proved extraordinarily difficult. Around that time, I became more convinced by statistical learning methods, though their primary application then was natural language processing.

In 2005, I was working on generative models for writing couplets. One day, something felt off — my program could compose proper couplets, yet it seemed to lack genuine mind.

When a person says "蓝田日暖玉生烟" (warm sun over blue fields, jade mists rising), the word "sun" evokes imagery, warmth, imagination. This is the connection between senses and language. I found this connection missing in the symbolic world. This led me to a perspective: intelligence should be an abstraction emerging from interaction with the physical world, which redirected my interest toward computer vision — a more foundational layer of physical world understanding.

The times moved incredibly fast. By 2008, I had the fortune of participating in ImageNet as a student leader. We discovered that computer vision, once considered an immensely complex problem, now appeared solvable.

Computer vision became my doctoral focus, similar to Yang Gao, so we actually met many years ago. After making certain progress in computer vision, we naturally wondered: could we return to the AGI question? Where does computer vision fit in the AGI framework? How far are we from achieving AGI? Vision is essentially the robot's eye — the conversion of physical signals into symbols, or internal representations for understanding and reasoning. Pushing spatial understanding further became part of my subsequent work: spatial intelligence research.

After what seemed like breakthroughs in spatial intelligence, we began seriously considering how to build general intelligence. This was around 2017. So my central question became: how should general intelligence even be defined? I felt we needed to think by analogy to humans, who possess numerous cognitive capabilities. If we had all of them, we might have an intelligence similar to humans.

At the time, I saw computer vision as mapping signals to symbols. But if a machine were to have self-awareness and deeper cognitive abilities, what was it most lacking?

I thought for a long time, and arrived at "a lack of conceptual emergence capability" — the ability to cognize whatever it's asked to cognize. It doesn't know what it is, nor what it should cognize. So the problem became: how to solve conceptual emergence. This line of thinking led me to define "embodied intelligence."

As humans, we speak of the unity of knowledge and action, of knowing and transforming the world as inseparable. Without "action," you cannot verify "knowledge." Without the goal of changing the world, there's no drive to understand it. But to change the world without form, without a body — that's impossible.

So we must consider achieving general intelligence with a form, with a body加持. From this perspective, I believe without embodied intelligence, there is no general physical intelligence, no general intelligence. Thus embodied intelligence becomes an outlet for general intelligence — this formed my understanding. I saw similar consensus among other scholars, so we began pushing for embodied intelligence. This understanding also shaped our choices in skill stacks and technical stacks.

Around 2022, through the success of ChatGPT and Transformer, we saw something interesting: regarding conceptual emergence, the Attention mechanism in Transformer enables certain compositional arrangements in concept space, representing a massive breakthrough in cognitive capability. Meanwhile, reinforcement learning and world models had also advanced significantly in our research — Professor Gao's earlier work had even inspired me somewhat.

Several developments made me feel embodied intelligence as a discipline had become relatively clear — the underlying logic was there, the core technical breakthroughs needed were identifiable. We further recognized that fully realizing such a framework for general physical intelligence, for embodied intelligence, would require enormous forces. Without capital support, this would be impossible to achieve within a university.

From both macro and technical understanding perspectives, it felt like the right thing to do. So we did it.

Oasis Capital: Thank you both. There's much debate now — some approach AGI from pure AI, others insist a body is essential. In your view, is the body strictly necessary? Or can we reach AGI directly through AI and multimodal approaches?

Professor suho: My entire intellectual starting point is: it must be.

I thought Jin Jian's framing was particularly excellent and inspiring for me. We are entering an agentic era. In this era, if it's a virtual agent, I think it will overall serve as an accelerator within commercial society and industrial workflows — an efficiency enhancement, a resource reallocation effect.

But if we push further on two questions — scientific discovery, for instance, truly requires unifying observation and experiment, which demands physical manipulation and experimentation. Or expanding human habitation to deserts, to deep oceans, establishing new cognition — do we need the unity of knowledge and action, the coordination of perception and interaction in new environments? Must humanity leave Earth? I believe so, and it won't take that long.

Following the earlier logic, with science advancing so rapidly and productive forces being unleashed, we can't remain in zero-sum competition on Earth forever. We're making preparations, and once ready, we'll leave — whether sending human bodies to Mars or terraforming the Moon?

From this perspective, I believe a broader AI vision, general intelligence, even superhuman intelligence, must pass through embodiment.

Professor Yang Gao: I strongly agree.

If we're truly heading toward AGI, we must have a body. Professor suho spoke from a deeply philosophical level; let me offer some concrete examples.

Take leaving a room: walk to the door, turn the handle, push it open. For humans, this is effortless, done unconsciously. But when training robots, you discover countless failure modes, many unimaginable.

For instance: not turning fully before pulling back; or turning fully but continuing downward until the handle breaks, without ever pulling back. Why? Under current simpler training paradigms, it lacks genuine interactive experience with the environment, so it cannot learn "I need to turn the handle to this position before I can push the door open."

This applies not just to robots, but to current large language models and vision-language models as well. This is why we've recently entered the post-training, RLHF era for LLMs — we need agents to interact with reality. Through these interactions, agents learn what's right and wrong, learning from experiential engagement.

I gave an embodied example, but the same holds for agents: booking flights, writing code. Everyone is doing this now, conducting reinforcement learning in Docker or virtual machine environments to strengthen these skills.

Even for non-embodied applications like VLMs, because our current dominant training paradigm is non-reinforcement learning — passive learning from internet image-text data, even the most advanced models like GPT-4, Gemini 2.5 Pro, often fail when asked "is this object to the left or right of that one in this image?" This knowledge is rarely mentioned in human language; humans don't need to say it — they see it. So there's no linguistic supervision teaching VLMs this knowledge.

All these examples point to the same conclusion: for virtual agents or physical robots to achieve general capabilities, interaction with the real world or environment is an essential component.

My conclusion: if we truly want AGI, truly want robots to handle physical world problems, interaction with real environments and experiential learning are indispensable — and embodied intelligence is indeed needed to achieve this.

Oasis Capital: Another question: since "having a body" is so important, a necessary path to future AGI, is there a technical gap between China and the US now? And where do you see China's advantages, in what dimensions might it prevail?

Professor Yang Gao: Fundamentally, China and the US are the two leading players, with no particularly major technical differences.

I've been to the US many times recently, and looking at this objectively: America's advantage lies in having more talent in absolute numbers, even if the top companies are comparable. This stems from America's longer history in embodied intelligence research and training systems. China's main embodied intelligence force was largely trained in the US as well.

Of course, at the top-level strategic design, our understanding is basically aligned with America's.

That's one American advantage. Another is that American capital is more willing to make "world number one" scale investments. This partly reflects different capital environments. I think China's major advantage lies in hardware-software integration.

I hadn't been to the US for a while, and when I spoke with NVIDIA labs and Stanford labs again, their daily concern was: "If my robot breaks and takes three months to repair, what do I do? The robot absolutely cannot break." Our repair cycle is at most one to two weeks.

Anyone building robots knows how painful repairs are, and whether a project succeeds depends on iteration cycles. In America, iteration cycles are dominated by hardware damage and repair time; domestically, we can average 3-10x faster. In America, hardware parts are mostly Taobao-shipped, taking at least a week at fastest. In China, if urgent, you can get parts tomorrow. This advantage extends beyond using mature robot platforms to include iterating and testing new hardware ideas in reality.

And in this embodied intelligence era, hardware-software co-design is extremely important. Robot design is like creating Adam and Eve — you start with what seems like good design, but discover massive limitations when applied to real manipulation tasks. You need repeated hardware redesign, including data collection, requiring extensive hardware co-design. That's point one.

Point two: embodied intelligence differs from LLMs. We often discuss scaling for LLMs; the bottleneck is electricity. After visiting the US, everyone was saying power plants need 2-3 years to catch up with LLM training's power consumption growth. But embodied intelligence has different bottleneck segments. When evaluating an industry, look at its shortest plank. For embodied intelligence, the current shortest plank is data. I'm a firm real-world data believer — Professor suho may differ, exploring simulators more.

For real-world data, China has massive advantages: lower labor costs, stronger organizational capabilities, and faster data iteration efficiency than America.

Objectively, both sides have advantages and disadvantages. As a company, we try to maximize our advantages while addressing our weaknesses.

Professor suho: Professor Gao's assessment is quite comprehensive, so I'll just add some supplementary points.

China's societal investment in embodied intelligence, I believe, exceeds America's. Because whether markets, enterprises, or capital, America has inertia — inertia toward succeeding in the virtual world. When it comes to physical directions, there's often little confidence in competing with China.

Second, supply chain issues. Professor Gao mentioned China's supply chain advantages; I'll add that America's current tariff policy uncertainty may significantly impact startups. Can you use Chinese supply chains? What tariffs apply? Do you need backup arrangements? Constantly worrying about this during very early technical iteration cycles is quite unsuitable.

Additionally, China has limited retreat options on embodied intelligence. Under current dynamics, winning in virtual AI against America would be difficult — not impossible, but difficult. Thus embodied intelligence becomes something that must be continuously supported. From government to enterprises to civil society, there's generally less room to retreat, we must push forward. That's my feeling, though the judgment may not be correct.

Finally, chips are indeed an issue for China, including edge chips and training chips. Especially edge chips — one hopes the US government handles this reasonably. Otherwise, we're counting on China developing sufficiently strong domestic chip supply chains.

Oasis Capital: Professor suho mentioned China's determination to win in embodied intelligence. In this context, many say there's excessive government and various participation, creating industry bubbles.

But as frontline entrepreneurs, you must see things we don't. What do you see that drives such commitment? And from an outside perspective, what should we expect as the next milestone or breakthrough?

Professor suho: My current stage is quite interesting — there is indeed significant bubble, but progress is also real.

The normal industrialization sequence should be: scientific framework clear, technical breakthrough achieved, then industrialization begins, exploring business models and scaling, right?

Embodied intelligence — not just embodied, all of AI — is currently in a state of "everyone doing research," from primary to secondary markets, everyone putting money in. Why? Because the problem is fundamentally so large, and the larger the problem, the earlier entry is required if it might work. By the time it's clearly viable, it's too late for you.

So my basic view: this is essentially "mass research." But because the pace is genuinely fast, the delay between research and commercialization is small. So on one hand you see bubbles, on the other you must recognize that across AI research, embodied intelligence represents the next breakthrough point, the outlet for general intelligence.

Ask computer vision researchers what they see? They see strong academic consensus, people testing and trying to identify scaling laws. Various data modalities — simulation, real robot collection, video — with accompanying algorithms, VLA, reinforcement learning, world models — you can see very real progress. So I think this is how to grasp the overall picture.

Oasis Capital: The next question: do you think world models are the solution? Recently even Yann LeCun is pursuing world models, making this direction quite hot.

Professor suho: The term "world model" has been somewhat abused. Or rather, not everyone means the same thing by it.

However, one point is shared: we need understanding of physical world commonsense. As I mentioned, converting physical signals into robot-usable representations — everyone agrees this matters. Video prediction, simulation, and other means can attempt to build such representations. I don't think it's a question of whether it's "the solution" — it's an important component within a relatively orthodox framework. Any attempt to solve problems without doing this is unreasonable.

Briefly on VLA. What does VLA claim? Vision in, language in, action out — this isn't wrong. But what's the specific architecture? A VLA not supported by world models — I believe this cannot be correct.

Oasis Capital: How far do you think the next breakthrough is?

Professor suho: It mainly depends on how you define breakthrough.

General intelligence is a relatively continuous progression. When Sam Altman said GPT-3, scaling law would achieve AGI in the long run — I suspect not just me, but everyone in investment, heard this as mere rhetoric.

But was it a breakthrough? Absolutely. AGI isn't achieved by one breakthrough; it requires rounds of breakthroughs, several breakthroughs. So the next breakthrough depends on whose perspective: for researchers, a breakthrough maintaining investor interest and confidence — I think that appears on a one-to-two-year timescale. For a breakthrough comprehensible to the general public, somewhat longer.

Oasis Capital: Professor Gao, since you primarily work on the VLA path — I recall in late 2023, when you told me about this path, I hadn't heard others discuss it. Now it's quite established terminology with some consensus. How do you view progress on this path, and what's your expected next milestone?

Professor Yang Gao: In late 2023, I only intuitively felt that logically, from first principles, there was no problem. But I couldn't predict whether the curvature would be 1.0, 2.0, or 3.0. But from current model capabilities, the evolution speed exceeds my expectations.

Regarding whether it's a bubble, how much real progress we have, and what the next breakthrough is — I think the best anchor is returning to the LLM discourse. Because LLMs in this domain have essentially completely solved our data problem.

Embodied models are, at least in early stages, a copy process — we're copying the successful experiences from LLMs' early stages into embodied models. Currently, we're solving a pre-training problem, giving models substantial prior knowledge. So the next breakthrough people may see is: at this point in time, models can sort of understand your instructions, execute somewhat reasonably, but not yet precisely.

As Professor suho noted, this is a continuous process, not necessarily discontinuous. Within this progression, I think what we'll see next is models becoming increasingly obedient, better completing more diverse physical behaviors, until one day you suddenly realize you can chain all behaviors together to complete a complex task — that would be the GPT-4 moment for embodied robot models.

I predict the "4" moment is harder to call, but I think the "3.5" moment is roughly 2-3 years away.

Oasis Capital: Final question: what distinctive value has Oasis brought you?

Professor Yang Gao: Oasis was one of our first-round investors, and our company history isn't long — perhaps a year and a half. Throughout our entire history, Oasis has truly mentored my thinking and growth around entrepreneurship. Jin Jian, Ivy, and I communicate frequently; they often ask high-level questions I wouldn't normally consider. Each time I think about these questions, the answer differs, but each answer becomes slightly clearer than before. I deeply appreciate these discussions.

Another aspect, including today's occasion, is something I cherish in my entrepreneurial journey. I can meet more founders, hear their stories, their experiences and lessons. This is extremely rare in the very long, very sparsely-rewarded process of entrepreneurship.

I was going to mention another Silicon Valley advantage: the density of entrepreneurial atmosphere. Why? I once went to a restaurant in Palo Alto, first went to the wrong one, but in that wrong restaurant discovered Scale AI's founder dining, then saw many familiar faces. In that atmosphere, many conversations are happening that are truly inspiring.

Under China's conditions, the density isn't as high. But Oasis gives me many opportunities for such conversations to start happening. So in both these aspects, I've felt very positively and grown tremendously. I hope for more such activities in the future.

Professor suho: In my professor days, my general expectation of investors was "give me money, leave me alone." But my feeling now, especially with Jin Jian and Ivy, is they've given me a sense that I can be friends with investors — something I found difficult to establish. Because my original mentality was truly "give me money," but now when I have problems, I genuinely ask them, genuinely seek help. And I've felt my questions aren't always easy to answer — they need to help me research materials, or connect me with people, investing considerable time, genuinely helping, genuinely reaching conclusions, even persistently "answering ten for every one asked."

This has deeply changed my view of the relationship between founders and investors — I think achieving this is quite remarkable, truly.

Oasis Capital: Thank you both, and thank you all for your time.