Spirit AI's Yang Gao: Scientists Aren't the Most "Reliable" Entrepreneurs, But Building a Startup Is Like Playing a Game | Oasis Vitality
Counselor on Vitality

Oasis Capital was the seed-round investor in Spirit AI, and has continued to increase its stake in subsequent rounds.
We see them starting from a software-hardware integrated foundation, deeply fusing cutting-edge algorithms with hardware platforms, aiming straight for the endgame of scaled embodied intelligence deployment.
This article is an exclusive interview by Intelligent Emergence, documenting Spirit AI co-founder Gao Yang's deep reflections on technical pathways, commercial choices, and the future of the industry. Enjoy.

Whether at the recently concluded WAIC (World Artificial Intelligence Conference) or the currently opening WRC (World Robot Conference), how can you assess a robot's true capabilities at an exhibition?
Gao Yang, co-founder of embodied intelligence company Spirit AI, offers a few tips:
For robots claimed to be able to fold clothes, try crumpling the clothes into a ball and tossing them randomly on the table — see if it can still complete the task. Or give it pants, jackets, and see if it can generalize across categories.
When the robot is operating, observe whether its movements are smooth and fluid rather than jerky and stuttered — this represents the coordination between cognition and action.
...
Gao Yang, who offered us these guidelines, is currently one of the most sought-after entrepreneurs in the embodied intelligence field — after completing his PhD at UC Berkeley, he chose to return to China and become an assistant professor at the Institute for Interdisciplinary Information Sciences, Tsinghua University.
In 2023, he founded embodied intelligence company Spirit AI together with Han Fengtao, former CTO of珞石机器人 (Robostar) — Han brings extensive hardware experience, having previously overseen mass production and shipment of tens of thousands of robots, while Gao has a research foundation in AI. This pairing of academic and industrial expertise has made Spirit AI one of the standout companies in this wave of embodied intelligence.
In the 19 months since its founding, they have raised over 1 billion RMB in cumulative funding. Their investor roster includes Oasis Capital, Huawei Hubble, JD.com, CATL, Shunwei Capital, and others.
Transitioning from the "ivory tower" of academia into the business world, Gao Yang must also confront stereotypes and biases against "scientist entrepreneurs" — but he doesn't shy away from them.
"Scientist entrepreneurship is, to some extent, not very reliable," he says. In his view, scientists pursue truth through interest-driven work, while entrepreneurship aims at commercial success. "I'm constantly acknowledging my own limitations. I know what I'm not good at, and I try to compensate."
Gao Yang likens entrepreneurship to "a kind of game," where interactions with investors and customers are all part of the leveling-up process. He's met with over a hundred investors. At first, his technical explanations were so abstruse that he "put people to sleep," but Gao Yang could rapidly adjust based on feedback. "Now dealing with investors has become much more natural — it's a growth process I enjoy."
In this young entrepreneur's office — his computer monitor still has a capybara plushie sticker — Gao Yang spoke with Intelligent Emergence about his journey from scientist to entrepreneur, as well as his views on technical pathways for embodied intelligence. Below is the interview transcript, lightly edited.


Intelligent Emergence: In robotics, your partnership with Mr. Han seems like a strong combination: one a scientist in software direction, the other an entrepreneur with substantial hardware experience. What were your criteria for choosing a co-founder?
Gao Yang: I thought about this for quite a while — how exactly should embodied intelligence be sold to customers? My conclusion today, which seems fairly obvious, is that you have to do software-hardware integration. You have to be the Apple of embodied intelligence, not Android.
Because in the early stages of technology, cross-platform capabilities are necessarily weak. Doing software and hardware well together is how countless industries began. Take personal computers, for example — initially, companies like IBM did both hardware and software. It was only after three or four decades that gradual specialization emerged.
I've done a lot of software myself, but basically no hardware. So I felt that having both hardware and software capabilities strong was particularly important for the first 30 years of this company.
On the other hand, many hardware people don't embrace change, or haven't recognized the shift. But Mr. Han recognized this change very early on — we were on the same wavelength.
Intelligent Emergence: What did you see in 2023 that gave you this idea of robotics entrepreneurship?
Gao Yang: Mainly seeing how ChatGPT transformed learning paradigms. Before ChatGPT came out, I myself didn't believe in what OpenAI was doing day in and day out. Many very senior professors at Berkeley thought it was nonsense. But once they produced GPT-3.5, we reflected and realized we'd been wrong. Following this logic, embodied intelligence is an inevitable phenomenon — it just needs some time.
Intelligent Emergence: You decided in 2023 that robots must be software-hardware integrated, yet some leading robotics companies still neglect the "brain." What do you think?
Gao Yang: Leading companies have their own logic. Their logic is that they're very good at hardware, and selling to education customers already keeps them quite comfortable — they can go public on this. Their optimal solution is to first secure the education market and not let others take it, since many other companies are now trying to enter this space. After going public, they can slowly expand into other areas. It's hard for one company to do many things simultaneously, especially when the education market already has fierce competition.
Intelligent Emergence: If I make non-humanoid hardware, a new form factor — is there room for a company that only does the platform to grow?
Gao Yang: Platform design is strongly correlated with AI requirements. Let me give an example: I build a platform where when I extend my arm, inverse kinematics fails, so I can't reach something on the table. This kind of problem is very common. Without joint development of hardware and AI, you wouldn't even recognize this problem.
Intelligent Emergence: Just looking at this industry, can the market accommodate a second company like this?
Gao Yang: I think it's very difficult. Going from scientist to manager is a "game."

Intelligent Emergence: When Professor Wu Yi encouraged you to return from Berkeley, you were already planning to start a business. I recall you once mentioned that you felt returning to do research would be more challenging?
Gao Yang: At the time, I just wanted to return to China to do research — there wasn't the kind of technological transformation opportunity that exists now. My other option was to become a research engineer at a major US tech company. But that path was essentially planned out for you: here's this little thing, just do it well.
Being a professor, though, is like starting a lab from nothing — no equipment, no people, building everything from scratch. It's a 0-to-1 challenge. So I started the company around the second half of 2023, about three years after returning to China.
Intelligent Emergence: I sense you don't just consider things from a research perspective — you seem to think from a business angle.
Gao Yang: I'm very interested in how to make technology usable for everyone, so I started thinking about the business level — how to make robots well, which led to the conclusion of software-hardware integration, and then to choosing who would co-found with me.
Intelligent Emergence: Why do you consider management a technology? Because technology tends to be rigid and rational, while management has some emotional components.
Gao Yang: Management isn't strictly a technology. It's probably an intermediate state between technology and art. But management is traceable — yet unlike science and engineering where you can just follow the procedure, it still requires some adaptability and improvisation.
Intelligent Emergence: You previously mentioned that scientist entrepreneurship isn't particularly reliable. When you put it into practice yourself, how do you supplement these additional capabilities?
Gao Yang: Let me first explain why it's not reliable. Scientists pursue truth; it's interest-driven work. But in entrepreneurship, the most important goal is to make a product. Often it's not about truth, but about how to serve customers well. Different customers may have very different requirements and metrics.
In this process, you use the company as a vehicle to achieve this goal, and there are many specialized skills involved — like how to build a team, how to cultivate the company as a growing organism.
I certainly can't say I'll 100% succeed. I can only say I'm constantly acknowledging my limitations. I know what I'm not good at, and I try to compensate.
Intelligent Emergence: Specifically for you personally, how did you complete the transformation from scientist identity to entrepreneur identity?
Gao Yang: I think it's about acknowledging my limitations, openly learning this entrepreneurship system, and using commercial success to drive everything — not just the pursuit of truth.
Intelligent Emergence: Do you enjoy this process?
Gao Yang: I think I quite enjoy it. It's a pretty interesting game with many lessons. One lesson was that when I first talked to investors, I leaned too factual — I was very precise, but people were very sleepy, very bored.
Then I realized I couldn't present it this way — I needed a more vivid, engaging way to explain it to them. There are many lessons like this.
Intelligent Emergence: You enjoy this process too?
Gao Yang: In the objective world, this is what I need to accomplish. As long as I want to do this well, I have to go through it.
Intelligent Emergence: How many investors have you met? Have you kept count?
Gao Yang: I haven't counted, but probably one or two hundred. And with each one, you have to give them the pitch.
Intelligent Emergence: In this process, how do you continuously refine your approach to interacting with investors?
Gao Yang: I think feedback is extremely important — otherwise you don't know what you're doing poorly. Now dealing with investors has become much more natural. It's a growth process I enjoy.
Intelligent Emergence: Do you think this represents a significant challenge for you?
Gao Yang: I think it's fine. It's probably like any other technology — just a somewhat special one.

Intelligent Emergence: At this stage, using Transformer for pre-training is already consensus, but in the later stages of engineering at various companies, will there be clear differences in outcomes?
Gao Yang: I think you can go to the WRC venue and see for yourself. There may be thousands of theories, but you have to experience it yourself. For example, can you interact with it? Crumple up clothes and throw them to it — see if the robot can refold them.

Intelligent Emergence: This could serve as a guide for evaluating robots at exhibitions.
Gao Yang: Because robotics is a very complex system, it's hard to figure out who's better. I think the best method is to experience it yourself — see what each company's model can actually do.
Intelligent Emergence: Everyone's talking about VLA (Vision Language Action) this year. How do you judge which VLAs are better?
Gao Yang: One aspect is algorithms. For example, some VLAs can't decompose tasks. Spirit AI's VLA has a fast-slow system that makes movements very smooth. Robots without this system will have stiff, jerky movements.
Another aspect is data. Large models need substantial data for training. Our own models use internet human video data for pre-training. Some VLAs can't pre-train on human videos, so their performance is comparatively weaker.
From a technical perspective, it's these two points: what characteristics the algorithm has, what data is used for training, and how the data is cleaned, processed, and proportioned — these all affect outcomes.
From an observational perspective, it's how complex a task the robot can handle. Some models can only do relatively simple tasks — what we call pick and place. But models like ours can do complex tasks like folding clothes. You can mess with it a bit, and it completes the task very well.
Intelligent Emergence: Is Spirit AI's Spirit v1 VLA model derived from your two original research projects (ViLa and CoPa)?
Gao Yang: Not just those two — it evolved from much research, including OneTwoVLA, which was engineered into Spirit AI's model.
Intelligent Emergence: What differentiates your OneTwoVLA from typical VLAs?
Gao Yang: If you tell it something slightly complex, like "put the phone in the drawer," it might take three steps — pick up the phone, open the drawer and put it in, then close it. Typical VLAs can't do this. OneTwoVLA can autonomously decide when to decompose tasks into smaller subtasks and complete them. But if you tell it a very simple task, it won't unnecessarily decompose further.
Intelligent Emergence: You previously made a prediction that in four years we'll reach a "Robot GPT-3.5" stage. What characteristics does this stage have?
Gao Yang: At the Robot GPT-3.5 stage, basically you can tell it anything and it can complete 70-80% of tasks. For example, entering a home: "Go get me a bottle of water from outside the door." But it may not work 100% of the time — maybe only 70%.
Intelligent Emergence: The industry has made many reflections on the VLA pathway recently. What aspects do you think are still improvable?
Gao Yang: I agree with what Chen Jianyu (founder of Star1 AI) said before — the "L" in VLA is indeed too much right now, because this model doesn't actually need to understand such complex language. VLA does have considerable room for technical improvement.
Intelligent Emergence: Specifically, how to improve?
Gao Yang: Getting down to brass tacks, there are many aspects. At the data level, for example, how to better utilize internet human video data. Because while robots already widely use internet image-text data, Spirit AI is already using internet human video data, since human videos are intuitively related to robot tasks.
Second, how to use teleoperation data for continuous effective supervised fine-tuning of VLA, and how to enable VLA to do reinforcement learning in the physical world? Because supervised fine-tuning involves humans collecting data for it, while reinforcement learning is the robot doing it itself.
At the architecture level, as Professor Chen mentioned, how to reduce the "L" further. Also, how to design better action tokenizers — these are all areas that can be continuously explored and improved.
Intelligent Emergence: Is the fast-slow system also a uniquely developed technical point? When was this completed?
Gao Yang: Yes, this was about 4 months ago.
Intelligent Emergence: After the fast-slow system was developed, what were the major improvements in terms of movement?
Gao Yang: You see some robots doing things — jerk, jerk — that's because the model lacks a fast-slow system.
With our model, when we fold clothes there's a step where we flick the garment, and this movement needs to be very fast. If it's not fast, the clothes won't flick up properly. If you pause, it loses momentum.

Intelligent Emergence: People are still discussing world models today. In Spirit AI's R&D roadmap, is this being considered?
Gao Yang: I think the cost of world models is indeed quite high. Currently, embodied intelligence doesn't urgently need world model training, but I think ultimately it will be necessary — it's an indispensable part of RL (reinforcement learning). At our current stage, we have some small-scale training and use of world models, but nothing particularly large-scale.
Intelligent Emergence: Do you think a hierarchical approach is viable?
Gao Yang: I think hierarchies will ultimately be eliminated. It's equivalent to using human wisdom to decompose tasks into smaller pieces. Hierarchical approaches may have decent short-term results on some tasks, but long-term they're definitely not scalable, because every new task requires manual work. But with end-to-end approaches, you just need to supplement data to the model.
Intelligent Emergence: In your view, what are the remaining non-consensus gaps in robotics?
Gao Yang: I have many closed-loop ideas in my own mind, but for example, the importance of end effectors, the first wave of robot landing scenarios — there are still many non-consensus areas. Including VLA algorithms, which are in a rapid development phase. The basic framework is set, but algorithmic details are still evolving quickly.

Intelligent Emergence: What do you think of the phenomenon of some robotics companies building large data collection factories? Could there be a problem where data collected by one company can't be used on another company's different hardware?
Gao Yang: I think large-scale data collection factories don't have much value at this stage. The main reason is that robot form factors are still constantly changing. When the robot form changes, previous data can't be 100% migrated — it gets significantly discounted.
On the other hand, according to our own algorithms, you don't actually need such large-scale data collection factories. I think the most important thing is to do pre-training well; data collection is secondary. I think there's a bit of putting the cart before the horse right now.
Intelligent Emergence: I sense some vendors are also treating this as a business model?
Gao Yang: I think it can indeed generate some commercial revenue in the short term. Many AI companies in the US — labor is too expensive there, so they can't build data collection factories, and they'll buy some data. But long-term, I think this model is hard to sustain, because the cross-platform problem hasn't been solved.
Intelligent Emergence: But the data they buy — if used on their own incompatible platforms, is this data still valuable?
Gao Yang: It has value, but at a discount.
Intelligent Emergence: It feels like robot demos are somewhat homogeneous now — why are they mostly folding clothes, opening appliance doors?
Gao Yang: First, folding clothes is widely recognized as the most difficult task, because clothing shapes vary endlessly and are very hard to pre-program. Actually, from demos you can see the differences in companies' model capabilities, so people like to do this.
Then, opening refrigerator and washing machine doors — these are everyday tasks that let people imagine the future.
Intelligent Emergence: What proportion of your data comes from the internet? What are the respective roles of different data types?
Gao Yang: By volume, over 95%. Internet data covers very broad scenarios and serves a pre-training role. Its main significance is providing data diversity — academically speaking, we want the model to generalize. The essence of generalization is that the robot has seen sufficiently diverse data.
Teleoperation connects generalization with precise physical-world manipulation. Because if the robot only watches others do things without doing them itself, it's hard to accomplish anything. Teleoperation provides precision.
Intelligent Emergence: How does generalization manifest?
Gao Yang: For example, the robot picks up my phone — mine is a foldable, but it was trained on iPhones. It can recognize the form factor and weight without needing foldable-specific training data.
Intelligent Emergence: How is generalization performance across the robotics field generally?
Gao Yang: It's still at a relatively early stage. But we've found that after using internet data, robots' generalization improvement is quite substantial — for example, when you swap objects, there's 60-80% improvement. Ultimately, pre-training and teleoperation data mixed together help each other.

Intelligent Emergence: The "Berkeley Four" — you four have very similar research directions and backgrounds. What are the specific differences in research approaches?
Gao Yang: Professor Chen Jianyu's background is in MPC (Model Predictive Control). When he first returned to China, he worked on safety RL, which is control theory. He later started working on humanoid robots, focusing on walking and running.
I myself lean more toward manipulation — using robot hands to do tasks, in the imitation learning and RL fine-tuning framework.
Huazhe Xu mainly works more on 3D policy — for example, using point clouds for manipulation and recognition. His DP3, for instance, achieves manipulation by capturing scenes with 3D cameras.
Intelligent Emergence: Do you privately compare whose direction is closer to the endgame?
Gao Yang: Everyone freely chooses their research direction, and each person's thinking certainly differs somewhat. Academically, I think it's hard to convince each other.
Intelligent Emergence: Do you privately discuss management?
Gao Yang: When everyone first became professors or just started companies, we all faced learning processes in management — we discussed these a lot.
Intelligent Emergence: From a memorable discussion among you four, what management conclusion did you reach?
Gao Yang: I remember once discussing with Huazhe Xu how they recruit versus how we recruit, commiserating that finding really great people isn't easy. We also discussed how to interview people.
Intelligent Emergence: DeepSeek's hiring logic involves having many young roles in the team. Is your logic similar?
Gao Yang: LM and VLM, and robotics are somewhat different, but the basic profile may be similar — relatively young, relatively smart, perhaps not necessarily with extensive work experience. Actually, we don't need many people, but we need strong ones.
Intelligent Emergence: "Strong" — how do you define this?
Gao Yang: A typical profile would be someone with a master's or PhD from a good school. They may have published several papers in robotics but haven't worked in companies, yet already have research experience.
Intelligent Emergence: Why isn't strong work experience needed? Is it because your own experience working in companies wasn't great?
Gao Yang: Not at all — it's just that robotics technology is changing too fast. For algorithm positions, if someone has worked in a company for three to five years, they probably studied longer ago, and the technology then was completely different from now. Their education may not match what we currently need. We need young people because the technology they're exposed to right now is the most cutting-edge.
Intelligent Emergence: From your four backgrounds, you all migrated from autonomous driving. Looking at the big picture, what overlaps between autonomous driving and robotics, and what incremental parts do you need to add later?
Gao Yang: The overlap is that the essence of these two problems is similar — you see a scene, make an action, and this action is either the robot moving forward or grasping something, or the autonomous vehicle moving forward or braking.
But there are also many differences. For example, autonomous driving platforms are ready — you don't need to build them; there are twenty to thirty car companies that can build cars very well. But humanoid robot platforms are still in rapid development.
Also, autonomous driving has extremely high safety requirements, while humanoid robots, in certain scenarios, have relatively lower safety requirements — their error tolerance is much higher.





