"I Saw the Signal of Scaling Law" | A Conversation with Tsinghua University IIIS Assistant Professor Mengdi Xu on Embodied Artificial Intelligence, World Models, and True Generalization

👦🏻 Podcast Interview: Koji

🥷 Edited by: Crossing

🧑‍🎨 Layout: Zeoooo

🚥 This week's guest on Crossing is Mengdi Xu, an assistant professor at Tsinghua University's Institute for Interdisciplinary Information Sciences (IIIS). Starting from vehicle engineering at Tsinghua, with stops at Johns Hopkins, CMU, and postdoctoral work in Jiajun Wu and Fei-Fei Li's groups at Stanford, she returned to China in early 2024 to join Tsinghua IIIS. Her core research focus is In-Context Learning (ICL) for robots: getting a robot to enter a new environment, learn a new task through just one or two interactions on the spot, and get faster at learning as it goes.

This episode also has a timely postscript: the week after we recorded, Generalist released GEN-1.5 — by showing a robot just a 3-12 second demonstration as a "physical prompt," with no fine-tuning or parameter updates, it could attempt new tasks immediately (averaging about 59% success across 10 tasks). ICL instantly became one of the hottest topics in global embodied AI, and this is precisely the direction Mengdi has been betting on.

Beyond ICL, we also discussed:

  • The ceiling of pre-training
  • True generalization
  • The signals and traps of Scaling Laws
  • Building stable data flywheels
  • The misleading nature of folding-laundry demos
  • Industry bubbles vs. genuine progress
  • The real life and research choices of a young PI

Beyond all the technical topics, this remains a conversation about life choices, environment, and long-term thinking — we hope it can serve as a reference for all young people standing between China and the US, academia and industry, following trends and holding the long line.

Listen on WeChat:

Listen on Xiaoyuzhou:

🎬 The video podcast is now live on Koji's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms

Rapid-Fire Q&A

👦🏻 Koji

Let's start with the traditional rapid-fire round to help everyone get to know you. Mengdi, how old are you?

👩🏻 Mengdi Xu

I'm 30.5 years old.

👦🏻 Koji

30.5? So precise! MBTI and zodiac sign?

👩🏻 Mengdi Xu

My MBTI is INTJ, and I'm a Capricorn.

👦🏻 Koji

In one sentence, what's the research direction you currently spend the most time on?

👩🏻 Mengdi Xu

Getting robots to continuously improve themselves from real-world scenarios.

👦🏻 Koji

How long have you been working on this? I remember you were working on generalization quite early on.

👩🏻 Mengdi Xu

After CMU, I've basically been following the generalization thread.

At the time, I felt robots couldn't just rely on pre-learned knowledge. What makes humans so generalizable operates on two levels: one, we have lots of prior knowledge; two, we learn new things extremely fast. True generalization means giving robots this kind of fast-learning ability similar to humans.

👦🏻 Koji

What was the catalyst that got you into this research direction?

👩🏻 Mengdi Xu

After arriving at CMU, I started working in what was then called Robot Learning — applying machine learning methods to robotics, before it had this fancy name "embodied AI."

At CMU, I really loved a work called MAML, a famous result by Stanford professor Chelsea Finn, who is also co-founder of Physical Intelligence. What made MAML powerful was that it could get robots to learn tasks they had never seen. I felt at the time that this was the future direction for robotics.

👦🏻 Koji

How did you come across this work?

👩🏻 Mengdi Xu

This work was very well-known at the time. There were labmates very interested in cutting-edge ML directions, and we would frequently exchange ideas and share papers. We read very broadly, so we noticed this direction.

Hengshui, Tsinghua, JHU, CMU, Stanford

👦🏻 Koji

At Tsinghua you studied vehicle engineering for undergrad. From vehicle engineering to embodied AI — that feels like there's a Transformers gap in between. How did you make the transition?

👩🏻 Mengdi Xu

In undergrad I studied automotive, officially called vehicle engineering. Looking back, choosing the automotive department was absolutely the right call — the curriculum was very comprehensive: mechanical design, automotive electronics, strong cross-pollination with the electrical engineering department, plus lots of algorithm courses. From mechanical design to algorithms, I got exposure to everything as an undergrad.

The real pivot to robotics came in my second year. I was doing research with Professor Shaoze Yan from the mechanical engineering department, working on bio-robotics, studying the superelastic recovery of bee wings. During pollination, bees' wings frequently collide with their surroundings, but they don't get damaged because their structure has superelastic properties.

Following this thread, for my third-year summer research I went to Professor Howie Choset's group at CMU, which is very well known for bio-robotics. I designed a snake robot for pipe crawling, and from that point on, my work gradually shifted from mechanical design toward robotics algorithms.

👦🏻 Koji

This doesn't sound like the typical path for a vehicle engineering student?

👩🏻 Mengdi Xu

Automotive students have a particular trait: when they want to do something, they'll work extremely hard to find a place that matches their interests.

There were some seniors in Professor Yan's group at the time, including Heng Yang, who's now an assistant professor at Harvard, plus a high school classmate who had also done competitions with me. By a fortunate coincidence, I was very lucky to get into Professor Yan's group.

👦🏻 Koji

Speaking of high school competitions, what was your experience like in middle school and even elementary school? Were you a top student all the way through?

👩🏻 Mengdi Xu

I really don't dare call myself a top student. Later in my studies and abroad, I've met too many impressive people, each with very unique ideas. Whether someone is a "top student" or not just comes down to whether they channel their energy into academics or into something they're more interested in.

My elementary school years were quite relaxed — no cram school, just school activities or hanging out with friends in my free time.

👦🏻 Koji

I remember your high school was Hengshui High School?

👩🏻 Mengdi Xu

Yes, Hengshui. My family gave me a pretty relaxed environment in elementary school, so I had a happy childhood. I had a hobby of building four-wheel-drive cars — the anime Bakusō Kyōdai Let's & Go!! was hugely popular at the time, so I'd buy parts and build them myself. Building cars was expensive, so we figured out ways to make money.

👦🏻 Koji

You could make money playing with four-wheel-drive cars?

👩🏻 Mengdi Xu

Several kids in the compound were into it, so we found ways to raise money for cars.

👦🏻 Koji

How did you earn it?

👩🏻 Mengdi Xu

Picking up scrap on the way home from school and selling it — actually quite profitable.

👦🏻 Koji

Sounds like an urban treasure hunt, but what kind of scrap could sell for four-wheel-drive car money?

👩🏻 Mengdi Xu

Some discarded metal parts — there were people who would buy them. Thinking back, I had quite a bit of persistence as a kid searching for parts.

👦🏻 Koji

Sounds like you planted the seed for vehicle engineering pretty early?

👩🏻 Mengdi Xu

It was indeed one reason I later chose vehicle engineering. I've always found mechanical structures really cool — assembling many tiny parts together, condensing generations of craftsmen's accumulated knowledge into a precision instrument.

👦🏻 Koji

What was the Hengshui High School experience like?

👩🏻 Mengdi Xu

My personal experience was quite positive. In high school I was a competition student — my main focus was competitions, and the management requirements were different from the regular college-prep classes. Competitions require more self-driven motivation, so the schedule was different from regular classes.

👦🏻 Koji

How did you become a competition student?

👩🏻 Mengdi Xu

At the end of the first semester of freshman year, there would be an interest survey, and teachers would ask which subject you were interested in. At the time I was choosing between math and physics, and I felt physics was a very grounded, real-world discipline, so I chose physics.

👦🏻 Koji

I've heard Hengshui starts at 5 or 6 in the morning, with schedules precise down to the minute. I've even heard you only get one shower per week?

👩🏻 Mengdi Xu

Whether in college-prep or competition classes, the schedule requirements were very consistent: morning exercises, mid-morning exercises, extremely strict timetables, with different grades going to the cafeteria at staggered times. The management was indeed very rigorous.

👦🏻 Koji

There's a lot of outside discussion about the Hengshui model, and people are curious about the future development of Hengshui graduates. What's your observation?

In the competition classes, everyone had one defining trait: extremely strong individual opinions. Competitions demanded enormous time commitment — while others got winter and summer breaks, we were in intensive training camps. Some of my classmates went abroad and became excellent researchers; others stayed in China and have done very well.

👦🏻 Koji

Choosing the competition track back then — was it like choosing a path of no return?

👩🏻 Mengdi Xu

It felt that way. You couldn't study all subjects comprehensively like the regular gaokao students. You went through competition exams, direct admissions, independent recruitment — and only when保送 was completely off the table would you merge back into regular gaokao classes, by which point it was already the second semester of senior year.

👦🏻 Koji

A bit like entrepreneurship — burning your boats.

👩🏻 Mengdi Xu

Exactly, but at the time it felt incredibly interesting and energizing. The people around me were completely aligned in their goals, from solving a single problem on a small scale to winning provincial first prizes or gold medals on a large scale. The motivation was so strong that it never felt like suffering.

👦🏻 Koji

Was there a lot of competition among classmates?

👩🏻 Mengdi Xu

Competition definitely existed, but our coaches guided us well. In the classroom, we felt more like comrades fighting side by side. Everyone started from zero and huddled together to solve problems — this kind of camaraderie was far stronger than any competitive tension.

Pre-training Alone Isn't Enough

👦🏻 Koji

At CMU, studying mechanical engineering, you won the department's best PhD thesis award. How many does that go to each year?

👩🏻 Mengdi Xu

Two people in our cohort.

👦🏻 Koji

What was the core content of that thesis?

👩🏻 Mengdi Xu

The title was Building Adaptable Generalist Robots. The core goal was giving robots the ability to adapt to dynamically changing environments.

Pre-training is certainly useful, and I very much believe in scaling — but pre-training isn't everything. We can't expect robots to become 100% successful generalists through pre-training alone; they must adapt to specific changes when deployed in real environments.

Reality is an open world where tasks constantly change: a home might have 20 objects today and 30 completely different ones a week later; different people's preferences vary across different scenarios. Correspondingly, robots must be able to adapt to change and iteratively improve themselves within that change.

👦🏻 Koji

What were the main solution paths at the time?

👩🏻 Mengdi Xu

Mainly two:

The first was In-Context Learning; the second was Test-Time Training, adjusting a small portion of model parameters during deployment. Both aimed to free robots from pure reliance on pre-training, enabling continuous learning through context and interaction in real-world scenarios.

The In-Context Learning line originated from my 2021 work, Prompt Decision Transformer. The motivation came from seeing GPT-3's in-context learning capabilities. I thought it was so cool — traditional ML would fine-tune for new tasks, while GPT-3 just gave a prompt or output format and the model adjusted its output directly based on that prompt. This adaptation method was efficient and direct, no retraining needed on the deployment side. I wondered: could robots have this capability too?

Prompt Decision Transformer started from a relatively basic setting: give the model a segment of robot demonstration and have it quickly adapt to different locomotion and manipulation tasks.

The follow-up work was Hyper-Decision Transformer, which took the Test-Time Training route. Adjusting parameters during deployment is a double-edged sword: tweaking them solves the current task, but tweak too much and you forget pre-trained knowledge. We compromised by building a Hypernetwork that injected meta-knowledge, which then generated a small Adapter comprising only 0.5% of the total Transformer parameters. During deployment, only this 0.5% needed updating.

👦🏻 Koji

How did it perform in practice?

👩🏻 Mengdi Xu

With limited pre-training data, the Hypernetwork approach showed results more easily.

But as data volume and pre-training scale have grown massively, I believe the core direction going forward remains In-Context Learning. Solving through prompting without frequent parameter changes; or waiting until enough multi-task data is collected, then doing batch parameter updates.

👦🏻 Koji

How does your current work extend from what you did then?

👩🏻 Mengdi Xu

After Prompt Decision Transformer, much more industry work emerged using In-Context Learning for fast adaptation.

DeepMind had a work called Algorithm Distillation with a stunning perspective: what kind of intelligence should a model actually be equivalent to?

Current VLA approaches equate model parameters to carriers of task knowledge, distilling expert demonstrations into parameters through imitation learning; Algorithm Distillation argues that model parameters should be equivalent to an algorithm — an RL algorithm. It puts suboptimal trajectories into the context window, letting the model continuously correct future actions based on past mistakes and history. The Transformer structure thereby gains self-improvement capability; the parameters themselves become a reinforcement learning algorithm.

👦🏻 Koji

Has this been validated on real robots?

👩🏻 Mengdi Xu

It worked in simulation, but the tasks leaned more toward navigation and locomotion, like maze-running. Applying it to manipulation remains very difficult.

Similar work includes Skild team's LocoFormer, which uses Transformer-based In-Context Learning for legged robot adaptation — also locomotion-focused. And NVIDIA's recent RoboTTT, where robots observe long-horizon human manipulation videos and change behavior through TTT layers to reproduce demonstrated actions.

👦🏻 Koji

These solutions sound quite divergent — any signs of convergence?

👩🏻 Mengdi Xu

The industry's emphasis on adaptation has clearly risen.

People are gradually realizing that relying solely on massive data and parameter pre-training, or purely on expert demonstration-based imitation learning, cannot fully solve robot applications. Adaptation built on top of these has become the breakthrough point, and this consensus is forming.

👦🏻 Koji

This sounds like a natural, obvious thing — why has it only become core consensus now?

👩🏻 Mengdi Xu

There's been ongoing debate in the field. Human generalization ability stems from two sources: first, deep priors and cerebellar control from long evolution; second, the extremely fast speed at which humans learn new skills.

The previous mainstream approach was to scale data and models first, building thick priors for robots and strengthening multi-task and instruction-following capabilities — this didn't mean adaptation wasn't important.

But in real scenarios, if robots had only adaptive capability without a sufficiently solid base model, the trial-and-error cost would be terrifying — like smashing all the cups in a home. Building strong priors first, then doing adaptation: this path makes sense.

👦🏻 Koji

For your postdoc you went to Stanford, joining Jiajun Wu and Fei-Fei Li's lab — what directions did you focus on then?

👩🏻 Mengdi Xu

I worked on three projects I was very satisfied with and found deeply interesting. I had extremely strong collaborators, mainly pushing my PhD's more theoretical ideas toward real-world deployment.

The first was the BEHAVIOR Challenge. This was a simulation benchmark that Fei-Fei's group poured enormous effort into, trying to answer a fundamental question: what real-world tasks should robots actually solve? It contained extremely rich household and school scene assets. I was responsible for dataset generation, proposing a method called MoMaGen that generated bimanual mobile manipulation data — closer to practical deployment than my past work.

The second centered on making robots more steerable, part of my main research thread. Enabling robots to follow more diverse human instructions and adapt to different people's preferences.

The third was a benchmark called Creative Tool Use, exploring whether robots could possess the human ability to creatively use physical tools.

👦🏻 Koji

How do you evaluate something like that? From the name, it needs to be both creative and benchmarkable.

👩🏻 Mengdi Xu

We started by collecting real human data. When people use tools, they never look at the tool's name — they look at function, what in robotics we call affordance.

Want to cut cake but no knife? A fork or chopsticks works fine. Need to clear a table but no broom? A book or even an arm sweep gets it done. Humans creatively exploit objects in their environment; we wanted robots to have this capability too.

We collected large amounts of human creative tool-use data and evaluated across three dimensions: first, multimodal large models' ability to select tools and plan trajectories; second, 3D and 4D scene reconstruction capability for non-rigid, deformable objects; third, whether skills from humans dexterously using tools could transfer to robots.

Training Under Fei-Fei Li

👦🏻 Koji

What's Fei-Fei Li's style like, and what's the atmosphere in her group?

👩🏻 Mengdi Xu

I had two advisors at the time, Fei-Fei and Jiajun. I was incredibly fortunate to have two advisors — their styles were quite different, and I gained very different things from each.

Fei-Fei is very vision-driven. She always throws out questions that push you beyond any single project — like: What is the most fundamental question in this field over the next ten years? What role do you want to play in it?

👦🏻 Koji

Were those her exact words?

👩🏻 Mengdi Xu

Not exactly, but every deep conversation eventually pointed toward that level of thinking.

When we were setting benchmark metrics, we tried to define many fine-grained process indicators — progress percentages, time spent at each stage. Fei-Fei cut straight through it: which of these metrics actually matters? Only the final success rate. Process metrics are for debugging during development, not for the ultimate vision. She quoted Einstein at the time: everything should be made as simple as possible, but not simpler.

As a PI, the core responsibility is to keep everyone locked on the most essential goal. Jiajun, on the other hand, is very hands-on — from project selection to specific design, he deeply reasons through everything, and when I felt lost, he pushed me to refine my research brand. The two of them made an excellent co-advisor pair.

👦🏻 Koji

How do you understand the idea of building a research brand?

👩🏻 Mengdi Xu

It has to grow out of what you genuinely care about. If a problem is something you're willing to spend five to ten years digging into, it naturally becomes your research brand.

👦🏻 Koji

That's a lot like entrepreneurship — it has to start from something you truly believe in deep down?

👩🏻 Mengdi Xu

Yes, start from what you believe in.

👦🏻 Koji

That reminds me of Paul Graham's essay How to Do Great Work — every field looks like a perfect smooth sphere from afar, but up close it's full of pits and black holes, and those flaws are exactly where entrepreneurship or research begins.

Returning to Tsinghua

👦🏻 Koji

What later prompted your decision to return to China? Staying at Stanford seemed like the obvious path at the time.

👩🏻 Mengdi Xu

I interviewed at Tsinghua's Institute for Interdisciplinary Information Sciences (IIIS) in 2023. I talked with many faculty there, including Professor Yao.

👦🏻 Koji

How did you decide to go for that interview?

👩🏻 Mengdi Xu

In 2023 I entered my fifth year of PhD and was thinking about next steps. For me it was mainly a choice between industry and academia.

I had interned at Google for a very long time at that point, and the experience was excellent — my mentor gave me tremendous help, including building my own research taste. In industry you do have resources to do things at larger scale, while academia has certain resource constraints. I thought a lot about whether to go to industry or academia.

The opportunity to interview at IIIS came through meeting Yi Wu, who helped bridge the connection and got me the interview. Conversations with faculty and Professor Yao during the interview all went quite well.

IIIS has several strengths. First, the people are exceptional — the students are extremely strong. Many people working in robotics and AI in the U.S. know someone from Yao Class, and there are top researchers across many fields among Yao Class alumni.

Second, my research direction is robotics, and robotics is inherently a system. Doing only the brain or only the hardware feels incomplete to me. I needed to find the best place to build robots into a system that can actually land in the real world.综合考虑下来,叉院的机会是我的最优选择。

👦🏻 Koji

Between industry and academia, how did you ultimately decide? Many people face that same Crossing right now.

👩🏻 Mengdi Xu

Before deciding, including before the IIIS interview, I sought advice from many friends and mentors. I eventually realized that what truly drives the decision is understanding what you specifically want.

Industry offers substantial compute support that makes it easier to scale up small ideas, but there's also a loss of freedom. Freedom matters enormously to me — I want to maximize how much time goes toward problems I consider important, and academia is relatively better for that.

It comes down to what you genuinely feel matters in your heart.

👦🏻 Koji

You mentioned before that a certain professor asked you a question that left a strong impression and made you think highly of them?

👩🏻 Mengdi Xu

I've kept thinking about that question.

He asked: Why do you believe you'll definitely become a top researcher in your field in five years?

My answer at the time wasn't very good — I just said with complete confidence, "I know I can do it." But I've gradually come to think that if I find a topic I'm genuinely interested in — like making robots one day enter millions of households, while continuously improving themselves in real environments — and pursue that along the right path for five years, I believe I won't turn out badly.

👦🏻 Koji

You mean being extremely focused on one path?

👩🏻 Mengdi Xu

Yes, focused, truly thinking deeply about that problem.

👦🏻 Koji

What do you remember from your interview with Andrew Chi-Chih Yao? What did he ask?

👩🏻 Mengdi Xu

He covered both life and research. Professor Yao asked why I wanted to return to IIIS, and my answer was similar to what I just said — I felt IIIS was a good platform with the freedom to do what I wanted to do.

One piece of advice from Professor Yao benefited me enormously during my postdoc: focus your attention more narrowly. There aren't many truly important problems, and one's time is extremely limited. If you spread your time across too many things, the results won't be good. You must focus on what you are genuinely interested in.

👦🏻 Koji

Sounds like you're quite curious by nature?

👩🏻 Mengdi Xu

I'm curious about many things.

👦🏻 Koji

So you have to constantly remind yourself to focus?

👩🏻 Mengdi Xu

I've since found that different fields have interconnected elements. Robotics itself is an integrative discipline — some people work on hardware, others on algorithms, learning algorithms. If the goal is application and real-world deployment, you need some understanding of all of these.

In the AI era, robotics is no longer just about hardware — you need a relatively comprehensive, holistic knowledge system.

👦🏻 Koji

So how do you balance curiosity and focus? On one hand you need to focus on adaptation and generalization, on the other you need to understand the field's dynamics, and there are headline news items every day.

👩🏻 Mengdi Xu

I don't follow headline news as much now — I read more papers and first-hand sources.

Robotics is an extremely hot topic domestically and internationally lately, but we still need to spend a lot of time on our actual research. It's easy to get lost if you're constantly following news.

What truly matters is the work at hand — the things you believe will change the field in one year, two years, or even longer. You need to settle down and do that work.

👦🏻 Koji

Is adaptation and generalization for robots what you consider most important now?

👩🏻 Mengdi Xu

Yes, giving robots in-context learning capability — not being defined solely by data seen during pre-training, but able to take information gathered from the environment — failure cases, human demonstrations, robot demonstrations, body gestures, and correction instructions — and become an agent that can operate on its own, continuously learning new tasks in the scene.

Pitfalls and Luck

👦🏻 Koji

From the outside, your life seems to have gone very smoothly. How do you feel hearing that? Would you describe yourself that way?

👩🏻 Mengdi Xu

I've actually fallen into many pitfalls and gone through many turning points. Looking at my trajectory, I started in design, then gradually shifted toward learning, algorithms, and models — each transition had its rough patches.

But if I had to summarize my growth, it's more about luck. At every stage I've had wonderful friends, teachers, and mentors supporting me. My own word would be "lucky."

👦🏻 Koji

Do you think "luck" falls from the sky? Or do you believe there's a system or method to obtaining it?

👩🏻 Mengdi Xu

How a person grows is partly self-determined, partly the environment you choose — and the environment shapes you in return. The friends I've made have all been excellent, and having a good support network where people help each other matters a lot.

👦🏻 Koji

So how do you choose your environment?

👩🏻 Mengdi Xu

When I choose, I mainly look at whether I can maximize my impact in that environment, or maximize my interests and capabilities.

👦🏻 Koji

Haha, from your background I thought you were going to say "just pick the best one."

👩🏻 Mengdi Xu

"Best" is a general metric, but for personal development, you need to pick what's most suitable for you.

You need to see whether your capabilities multiply in a given environment — like clearing levels in a game, treating it as a game you keep advancing through.

👦🏻 Koji

Before entering an environment — whether IIIS, CMU, or Stanford — how do you learn as much as possible whether it's truly the best fit, the place that most amplifies your abilities?

👩🏻 Mengdi Xu

Information gathering is crucial. I'll first see if friends can connect me; if not, I'll send emails to connect directly.

I've found that many researchers are extremely warm-hearted — as long as you express your needs properly, people are willing to help each other.

👦🏻 Koji

That's quite hard for many introverts to break through. Any small tips to share?

👩🏻 Mengdi Xu

I'm actually quite introverted myself. But when it comes to my professional domain, I need to maximize my methods.

Even as an introvert, you can follow certain protocols — like having AI help draft an email. Write a first version yourself, then have it polished to be more appropriate. There are plenty of tools now; use tools correctly, and there's nothing you can't accomplish.

👦🏻 Koji

You must get many emails or cold calls from students or juniors now. What kinds of emails are you more willing to respond to?

👩🏻 Mengdi Xu

Some are clearly pure AI-generated — I don't respond to those.

The most important thing in an email is demonstrating sincerity. Sincerity shows through whether you actually understand the field and direction; it's testing whether you can connect your interests to actual work. Many people have interests but don't follow through. If you can clearly demonstrate interest and articulate what you're willing to work on, that's a sign of being thoughtful.

Thoughtful students — I'm generally willing to talk further with.

👦🏻 Koji

You mentioned facing several transitions in your career. What was your toughest crossroads moment?

👩🏻 Mengdi Xu

The toughest choice was definitely whether to return to China — which was also, in a sense, choosing between industry and academia.

I deliberated for a very long time. Most of my research career had been in the United States, so I worried about whether I could adapt to coming back. There would definitely be many challenges.

I especially love a concept in ML called "no free lunch." There's no such thing as a free lunch — every decision comes with its weaker side. After much consideration, what I valued most was freedom: the freedom to build robotics as an integrated system and actually deploy it. All things considered, returning to China for a faculty position was the optimal choice.

Is There a Bubble in Chinese Embodied AI?

👦🏻 Koji

Since returning over a year ago, you've probably met with numerous embodied AI researchers and entrepreneurs. What's your overall impression?

👩🏻 Mengdi Xu

I had some sense of it before returning and found it very exciting — people were energetically working on different layers of the robotics stack.

My current feeling is: there's been substantial real progress, but there's also some froth. These two aren't mutually exclusive.

The good part is that from materials and sensors to hardware bodies, data, and models, you can see domestic companies tackling key problems with significant resources invested. I'm very optimistic long-term — real progress is happening.

At the same time, some froth exists because embodied AI hasn't converged yet. Everyone is betting on directions they believe in. The focus keeps shifting: first it was data, then gradually world models, then back to data again. With so much flux, froth is inevitable — like when PID parameters aren't tuned properly and you get overshoot and oscillation.

👦🏻 Koji

What does overshoot refer to?

👩🏻 Mengdi Xu

Overshoot is when PID parameters aren't tuned well, causing oscillations.

👦🏻 Koji

Last time you mentioned that last year, many embodied AI practitioners were saying "data first," then in the first half of this year everyone was discussing world models, and recently the conversation has swung back to data. What's driving this shift?

👩🏻 Mengdi Xu

People's perspectives have shifted slightly, but it's not a complete pivot — everyone is continuously thinking about what truly matters in the industry.

Both issues are actually very important. They're not mutually exclusive paths but rather two different layers. Embodied data is extremely scarce; data determines what the model can see. World models, compared to VLA, represent a new route that focuses more on representation and understanding of the world, unlike VLA's task-oriented approach.

I personally lean toward the world model route. Because it's task-agnostic. Tasks too easily become an unscalable approach — there are too many to enumerate. If a world model can learn the patterns of how the world changes, that's a more fundamental capability.

👦🏻 Koji

In China and the United States, which companies or labs are you paying closest attention to?

👩🏻 Mengdi Xu

I follow quite a few, mainly those aligned with my research direction and beliefs.

In the US, I definitely follow Physical Intelligence, plus big companies like Google and NVIDIA, and startups like Generalist and Rhoda. Rhoda has been experimenting with In-Context Learning recently — an interesting approach.

Domestically, there are also companies whose work I appreciate. For example, LingBot Technology's recent LingBot-VLA 2.0 release follows the In-Context Learning route, which is very meaningful.

Signals of Scaling Law

👦🏻 Koji

Have you seen any signals of Scaling Law recently?

👩🏻 Mengdi Xu

I've seen some results from other companies. For instance, scaling pre-training data from 100,000 hours to 1 million hours yields decent performance on held-out, unseen test sets.

But what people are measuring now is mostly what's called "loss." One tricky thing in robotics is that loss and success rate aren't directly correlated. Very low loss doesn't necessarily mean high success rate — there can be lots of oscillation.

For example, picking up a cup might take 20 steps total. The first 17 steps involve no contact with the cup. Even if those first 17 steps fit perfectly, if the error is large at step 18 when contact occurs, the task easily fails. This temporal imbalance means loss and success rate are mostly not directly related.

If future Scaling Law demonstrations show that more data and larger models lead to gradually increasing success rates on unseen tasks — that would be truly meaningful, the Scaling Law I want to see.

Robotics' GPT Moment

👦🏻 Koji

Everyone loves to ask whether robotics will have its GPT Moment. Do you think it will?

👩🏻 Mengdi Xu

I think it will. I strongly believe in scaling.

By comparison, current mainstream embodied models still follow a route of "some pre-training, then fine-tuning for the target domain" — analogous to language models around the GPT-1 stage, requiring targeted post-training or fine-tuning to solve specific tasks.

What I truly hope to achieve is closer to the GPT-3 moment: through In-Context Learning and prompting, giving the model task-relevant information without fine-tuning. Just telling the model what to do through context, reducing the reliance on fine-tuning.

👦🏻 Koji

In watching for Scaling Law to bring about a GPT Moment, which benchmarks are particularly worth paying attention to?

👩🏻 Mengdi Xu

Compared to language models, embodied AI has a physical body, so evaluation can't look at success rate alone. Failure modes and risky situations need to be communicated to developers and users through predefined benchmarks about where the capabilities' boundaries lie. So benchmarks are even more important.

Many companies now use closed internal evaluation standards that outsiders can't access. Among open-source benchmarks, I'd still recommend BEHAVIOR — it features longer-horizon tasks with more diverse home environments and object arrangements. Though people criticize it for the Sim-to-Real gap, simulation itself is also an environment.

👦🏻 Koji

Has BEHAVIOR been continuously maintained and evolved since you worked on it?

👩🏻 Mengdi Xu

It's been continuously maintained, and usability has gradually improved.

👦🏻 Koji

NVIDIA's Jim Fan gave a talk earlier this year arguing that embodied AI development will strictly correspond to stages in language models. Do you agree?

👩🏻 Mengdi Xu

I generally agree. The known properties in language models — scaling, Scaling Law, foundation models, and In-Context Learning capabilities — will also appear in embodied AI.

But the emphases will differ. Robotics has its domain-specific uniqueness:

First, safety needs to be prioritized even more. Before deployment, something like stress testing is needed before putting it into real scenarios. This is more critical and urgent than for language models.

Second, robotics context is more diverse. Because it has a body, there's greater flexibility, and observing the surrounding environment means context contains more varied meanings.

Overall, robotics models will definitely be scaled up, but the priorities will come in a different order.

👦🏻 Koji

What about real-time performance?

👩🏻 Mengdi Xu

Real-time performance is indeed very important. You can't expect an embodied model to take 20 minutes to react — it needs to interact with the surrounding environment in real time.

How Will Robots Adapt to Every Home?

👦🏻 Koji

What's the most important thing you've been working on in the past month or two?

👩🏻 Mengdi Xu

Recently I've been pushing forward along the In-Context Learning route.

I want to explain what In-Context Learning actually means in embodied AI and what capabilities it can ultimately enable.

A simple example: many robot demos show folding clothes into squares, implicitly assuming everyone folds clothes that way at home. But that's not true. I personally use a capsule wardrobe — my clothes are rolled up. I want a robot to come into my home and have me tell it: look, this is how I roll my clothes. Another person might want clothes arranged by color. Some people don't fold at all — they just hang everything up.

What In-Context Learning aims to solve is giving models the ability to be defined by end users. Users can change the model's specific capabilities and behaviors according to their preferences.

Suppose a robot doesn't know how to cook a certain dish at first. I demonstrate it. After learning from 1,000 similar scenarios, it might initially take half an hour to adapt. Gradually it takes only 10 minutes or less. What we're training is the model's ability to quickly learn new tasks. Deployed across more homes and service scenarios, continuously reinforcing this learning ability, it will learn faster and faster.

I strongly believe In-Context Learning is ultimately how models will scale up.

👦🏻 Koji

Where do you think the most likely point of future failure lies?

👩🏻 Mengdi Xu

The hardest point is whether we can establish a stable data feedback loop for the model.

This approach requires integrating with real scenarios, deploying in end-user homes or service settings. You need suitable mentors to teach the robot, and after half an hour of teaching there needs to be obvious progress — only then will end users be willing to use it, and only then can the data loop start turning.

There's a chicken-and-egg problem here: the base model needs to be good enough for learning speed to be fast enough; but you also need to get into scenarios first for users to activate those base capabilities. Whether this data loop can run smoothly is the riskiest point.

👦🏻 Koji

What kind of data loop have you actually gotten working?

👩🏻 Mengdi Xu

Right now we're doing tabletop tasks — adjusting the robot's behavior based on human corrections or demos. Starting with tabletop.

👦🏻 Koji

On real hardware?

👩🏻 Mengdi Xu

All on bimanual setups, dual-arm scenarios, with dexterous hands too. What we ultimately want is full-body dexterous manipulation.

👦🏻 Koji

What's the relationship between In-Context Learning and offline imitation? Where do they differ and overlap?

👩🏻 Mengdi Xu

Offline imitation, by its accepted definition, means having expert data and learning from it via supervised learning.

It essentially comes back to the distinction between Transformer and Algorithm Distillation I mentioned earlier: imitation learning treats the model as a task memory module. After large-scale training, you hope it can solve unseen tasks through interpolative generalization — for example, seeing blocks in 10 positions on a table, then handling a new position within that range. But it doesn't actually learn new tasks; it's just memorizing motion knowledge.

What we want is for the model to see diverse types of data purposefully — suboptimal data, human body gestures, language correction instructions, human and robot demos, and so on. This rich offline data can unlock far more capabilities than pure expert demonstration imitation ever could.

👦🏻 Koji

Speaking of the Interdisciplinary Research Institute, people think of Huazhe Xu, Sunny, Jianyu Chen — these assistant professors who are very active in entrepreneurship. Do you all get together lately, have meetings?

👩🏻 Mengdi Xu

Privately we're all genuinely busy, no fixed time for everyone to sit down together, but we do communicate one-on-one.

Our broad vision is the same — getting robots to scale up and deploy in real applications where they're actually useful. But the specific bets differ. Some are more bullish on world models, some want to fuse VLA with world models. In the near term that's actually good, because the technical direction hasn't fully converged yet.

👦🏻 Koji

What topics do you all share interest in?

👩🏻 Mengdi Xu

The common thread is how to scale up models — including discussions around data-related research, and how to get models into human environments. I've chatted with Sunny about related topics before too.

👦🏻 Koji

On data — you mentioned earlier that the hardest part is defining requirements. Why is that difficult?

👩🏻 Mengdi Xu

Data definitions aren't pulled out of thin air; they're determined by model capabilities. Defining data means thinking through what capabilities the model actually needs, then reverse-engineering how to collect that data.

If you want smoother motions, you need smooth data. If you want failure recovery — the ability to learn from mistakes — your data needs to capture failures and corrections. Data definition should be driven by model capability definitions, except right now everyone's definition of model capabilities is still relatively fuzzy.

For example, should the base model have more CV-oriented capabilities like segmentation, tracking, depth and normal reconstruction? In-Context Learning is another new capability — getting the model to learn from gestures, body movements, language instructions, and other multimodal signals. Different model capabilities in turn redefine data requirements.

👦🏻 Koji

Among the mainstream data collection approaches, where do you place your own bets?

👩🏻 Mengdi Xu

I'm a fusionist — different data teaches models different things.

Teleoperation data is closest to the robot's own body, so it's best suited for fine-tuning or post-training stages where you want direct relevance to the target domain. At that level, teleoperation data is definitely optimal.

Simulation data easily generates visual generalization across positions, textures, colors — making models more robust to visual variation. It's more useful in pre-training or mid-training, earlier stages. And simulation data is relatively clean, so it can help unlock capabilities from large-scale video pre-trained models.

UMI data sits in between — more robotic than human data, but cheaper to collect.

The cheapest is Human Data. I personally love Human Data. Humans are extremely mature embodied agents, able to conveniently collect scene data for us in daily life. The low cost means you can massively deploy across different scenes. Robustness and understanding of scenes — that's where Human Data shines.

Some people think human and robot morphology are too different, but human itself is a morphology. With good morphology transfer methods, it's incredibly powerful.

And what we learn from Human Data should be scene dynamics and how objects change during task execution — that's not so strongly tied to specific morphology shapes. Task-relevant information can absolutely be learned from Human Data.

Essentially, you need to use the right data in the right place, correctly building the link between data and model capabilities.

👦🏻 Koji

Today, many students want to work on embodied AI. For young students just starting their research path, what advice do you have?

👩🏻 Mengdi Xu

First: read widely, absorb enough.

When I talk with some students, they seem interested in fancy concepts, but when you press them on what specifically interests them, it gets fuzzy. The root cause is they simply don't know enough. Books start thick, then get thin — you need to know enough first before you can peel away layers and understand what genuinely interests you. You have to read extensively, talk to lots of people.

Second: actually get your hands dirty. Many people are interested in embodied AI, and the sooner you cultivate the ability to translate interest into action, the better. Only by actually doing it can you understand where the real difficulties lie — for example, real robot deployment is extremely experience-intensive. Building some early accumulation helps.

👦🏻 Koji

Do you have must-ask questions when interviewing PhD candidates?

👩🏻 Mengdi Xu

I ask: What's the paper you're most interested in? Or framed differently: Which paper do you think is the best work?

This is a research taste or opinion question, quite difficult to answer. It presumes the student has some understanding of the field and knows where their interests lie.

👦🏻 Koji

What's the best answer you've heard?

👩🏻 Mengdi Xu

One student was very interested in diffusion, but he didn't cite Diffusion Policy — he cited the original diffusion model paper.

I felt this student had read extremely broadly, not limiting himself to robotics applications. He could trace back to find where a paper truly originated. This ability to dig to the root, to find the starting point, is excellent.

👦🏻 Koji

Do you take undergraduates into your research group?

👩🏻 Mengdi Xu

I do. We have some undergraduates who joined to work with me from the start, and you can see tremendous growth. From having unclear interests, to doing more hands-on work, to developing stronger opinions.

👦🏻 Koji

How do you select people?

👩🏻 Mengdi Xu

For undergraduates the bar is relatively open right now — I want to give people a platform to understand cutting-edge research.

We invite many undergraduates to listen in, in a tutorial format: the first 30 to 45 minutes covers the field's development from a history-of-science angle, or breaks a topic into different schools of thought; the final 15 to 30 minutes has students present their own projects. The main goal is giving students opportunities to learn.

👦🏻 Koji

How frequent are your group meetings?

👩🏻 Mengdi Xu

Originally planned for weekly, but in practice it's biweekly.

👦🏻 Koji

Who handles the tutorials you mentioned earlier?

👩🏻 Mengdi Xu

The students do, and they're quite impressive. I'll discuss topics with them in advance, walk them through the first pass to check if their organizational approach is sound, and gradually they can synthesize on their own. Their learning speed is very fast.

MBTI 👦🏻 Koji

I do get pretty strong INTJ vibes from you. Have you had any F moments?

👩🏻 Mengdi Xu

I'm actually a very empathetic person.

👦🏻 Koji

Where does that show?

👩🏻 Mengdi Xu

I'm acutely attuned to shifts in others' emotions. In a group, I'm the one who tends more to others' feelings.

Team cohesion is built by key people, and being more F in non-work moments helps close distance with students and collaborators. I don't want us to be merely work partners.

👦🏻 Koji

Outside of work, what do you talk about with your PhD students?

👩🏻 Mengdi Xu

I'm quite grateful to my students. Doing a PhD with a young assistant professor is itself a risky, very challenging choice.

During PhD interviews I tell them directly: coming here means building the lab, possibly doing work that looks like miscellaneous tasks. They've all been very willing to take on the challenge.

Building the initial lab is a bit like building a startup — not a one-person startup, but something my PhD students and I construct together.

Our trust has developed well. Everyone shares the desire to do things well, each contributing their abilities and time in the right places, so cohesion builds relatively easily.

👦🏻 Koji

Do you do team bonding activities?

👩🏻 Mengdi Xu

We do — just went karaoke together a couple days ago.

I've organized team-building twice myself, and students sometimes invite me when they organize their own. The lab atmosphere I want to build isn't just work or academic collaborators, but friends you can trust for the long term.

👦🏻 Koji

Let's end on something lighter — what AI products do you use day-to-day?

👩🏻 Mengdi Xu

I use ChatGPT quite a bit. Lately I've been alternating between ChatGPT and Gemini — they give pretty different perspectives on some things.

I used to use AI more for knowledge work, like summarizing content or doing research. Recently I've started having more non-work conversations with it, discussing my take on events.

One fun thing lately was naming the lab. I gave AI some keywords and an initial list, and had it help polish or propose new ones. The names it came up with were actually pretty good.

👦🏻 Koji

Have you settled on a lab name?

👩🏻 Mengdi Xu

Not yet — still needs more thought.

👦🏻 Koji

A lot of people have mentioned an explosion of papers in their field lately. Have you felt that?

👩🏻 Mengdi Xu

Yeah, robotics papers have basically been doubling recently, and a lot of them are AI-written.

👦🏻 Koji

How do you deal with the changes in the research environment right now?

👩🏻 Mengdi Xu

I consume papers in a few different ways. Every day I check what researchers I follow are promoting — the author list acts as a filter, showing which papers they're actually proud of. In embodied AI, people are increasingly used to promoting their work on social media, including Chinese researchers on various platforms, so I skim through those daily.

You can clearly tell when a paper is heavily AI-written. Right now AI-written papers might pass peer review and get into conferences, but the overall quality isn't higher than human-written ones. They look substantial at first glance, but the arguments are full of holes and don't actually contain much. If I spot an AI-written paper, my interest drops.

👦🏻 Koji

How many kilometers into the embodied marathon?

Last question: people love using the marathon metaphor for embodied AI — 42 kilometers total. How far along do you think we are?

👩🏻 Mengdi Xu

I'd say around 10 kilometers — roughly a quarter of the way.

What really matters is getting robots into real-world applications. We haven't yet seen intelligent robots with strong generalization and versatility actually deployed at scale, so there's still a long way to go.

But I'm generally optimistic. In the next three to five years, I think we'll see a clear paradigm shift or inflection point driven by scaling. The next three to five years will be a really critical period.

👦🏻 Koji

You have to be optimistic to work in embodied AI or AI in general — pessimism is easy to be right about, but only optimism can lead to success. Thanks Mengdi, really enjoyed chatting today!

👩🏻 Mengdi Xu

Thanks Koji!

Join our membership group