So He Turned Toward Embodied AI | A Conversation with Wang Jiawei: Simple AI's Gen Z Chief Scientist
The rougher the seas, the pricier the fish.

👦🏻 Podcast Interview: Koji
🥷 Edited by: Crossing
🧑🎨 Layout: Zeoooo

🚥 This week's guest on Crossing is Jiawei Wang, 24-year-old Chief Scientist at Simple AI. USTC's gifted youth program, MSRA, DeepSeek, ByteDance Seed — this path led straight to the cutting edge of large models, but upon graduation, he turned toward embodied AI. Why?
What he wanted to leave wasn't large models themselves, but the state of simply following predetermined plans. As he puts it: "The bigger the waves, the pricier the fish" — embodied AI still lacks consensus on routes, evaluation, and products. The risk is higher, but it also gives young researchers room to define the problems.
In this episode, we start from GPT-6 Astra and discuss whether "OpenAI will demolish embodied models with dimensional reduction." Wang believes general-purpose large models will take on more and more understanding and planning, but when a cup starts to slip, robots still need action models that can react fast.
That's why Simple AI has gone all the way from HiFi-UMI data and embodied foundation models to Agentic OS to the physical hardware itself: open-sourcing 2,000 hours of data while accumulating tens of thousands internally; the model has already shown certain zero-shot generalization; and Agentic OS connects high-level intent with low-level action.
Finally, we discussed the latest work from embodied AI peers in the United States: GEN-1.5's validation of pretraining and natural emergence, π0.7's reinforcement of steerability, and whether Skild AI's in-context learning trained specifically for demos is overhyped.
Listen on WeChat:
Listen on Xiaoyuzhou:

🎬 The video podcast is also available on Koji's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms



🚥
Rapid-Fire Q&A
👦🏻 Koji
Let's do rapid-fire first. Your age?
🧑🏻💻 Jiawei Wang
24, turning 25 next month.
👦🏻 Koji
Educational background?
🧑🏻💻 Jiawei Wang
Elementary and middle school in a small county, high school at Hefei No. 1 High School. Took the gaokao in my second year of high school and got into USTC's gifted youth program. In 2020, I got a joint PhD opportunity between USTC and MSRA, and just graduated in 2026.
👦🏻 Koji
MBTI and zodiac sign?
🧑🏻💻 Jiawei Wang
ENTJ, though the E is probably only like 50%. Libra.
👦🏻 Koji
One sentence to describe Simple AI.
🧑🏻💻 Jiawei Wang
Simple AI is a general-purpose embodied intelligence company for the home. Our slogan is "Keep The World Simple" — make the world simpler, which is where the English name Simple AI comes from.
👦🏻 Koji
In the last podcast, I talked with Mengdi Xu, an assistant professor at Tsinghua's IIIS who did a postdoc in Fei-Fei Li's group. She said Fei-Fei loves to quote Albert Einstein: make things as simple as possible, but not simpler. What do you think?
🧑🏻💻 Jiawei Wang
That makes a lot of sense. There's a famous Occam's Razor principle in AI — over-engineering things often doesn't work. Keeping it simple and doing thorough experimental validation is how you prove something actually works.
A lot of past academic research, like in NLP and early CV, over-engineered model architectures. Then suddenly people realized that GPT, with an extremely simple architecture, just solved everything.
👦🏻 Koji
A few days ago I met a researcher whose team was doing similar training after OpenAI o1 came out. Since he works on infrastructure, he designed a very complex architecture. Then DeepSeek-R1 was released, and his team spent two weeks deleting code — removed 90% of it — and realized they never needed such a complex design.
🧑🏻💻 Jiawei Wang
Exactly.
👦🏻 Koji
What was your work experience before Simple AI?
🧑🏻💻 Jiawei Wang
Just finished my PhD, so mainly during my doctoral studies. Spent the longest time at MSRA, working with my advisor Qiang Huo on document intelligence research.
In 2024, when large language models started showing capabilities that small models couldn't achieve, I interned at DeepSeek's multimodal team and worked on pretraining data for DeepSeek-V3.
Later I felt language models were evolving from Chat toward Agent. Chat is hard to generate major commercial value, but Agent can truly improve work efficiency. So in 2025, I interned at an early Agent team inside ByteDance Seed.
👦🏻 Koji
Now you're Chief Scientist at Simple AI. Did you dream of becoming a scientist when you were a kid?
🧑🏻💻 Jiawei Wang
I actually did — I made that wish when blowing out birthday candles.
👦🏻 Koji
What's the same and different between the scientist you imagined as a kid and what you actually do today?
🧑🏻💻 Jiawei Wang
Back then I imagined doing biology or chemistry experiments in a lab, or proposing conjectures in fundamental math. I had zero concept of what a computer scientist was — that's the biggest difference.
The similarity is that I do get to explore areas where nobody knows the answers yet.
👦🏻 Koji
What research topic do you spend the most time on right now?
🧑🏻💻 Jiawei Wang
How to properly do pretraining for embodied foundation models. Currently the various technical routes haven't converged at all — we're carving our own path.
Can GPT-6 Become a Robot Brain?
👦🏻 Koji
A couple days ago, GPT-6 Astra was released, and social media was flooded with videos of it directly driving robotic arms to do various tasks. What was your reaction seeing those demos?
🧑🏻💻 Jiawei Wang
Pretty shocking. At first it was hard to imagine that a general-purpose language model could achieve such high success rates on embodied tasks like grasping blocks — the capabilities exceeded expectations.
Actually, using LLMs for robotics isn't a new concept. The industry has been working on Code as Policies for years, and a few months ago Fei-Fei Li's team and Jim Fan's team released CaP-X. But early Code as Policies still relied on small models for perception and recognition. GPT-6 natively fused in the perception and localization capabilities of small models, fully leveraging 3D spatial understanding, scene cognition, long-horizon planning, and trial-and-error reflection.
One demo really stuck with me: a pen was strapped to the gripper to draw the Golden Gate Bridge. Looking at the finished product, it's impressive. But looking closely at the process, it actually drew it four times. The first pass was clearly figuring out how to use the pen. The second pass sketched the outline. The third and fourth passes is where the details actually became clear. Much of the capability comes from task trial-and-error, understanding, and experience summarization, combined with underlying localization and spatial understanding.
👦🏻 Koji
When you say four times — after each attempt, a camera feeds back and it adjusts itself?
🧑🏻💻 Jiawei Wang
It reads the video stream in real time. If one drawing isn't good, it discards it and switches to a fresh sheet of paper, but all the previous context stays in the context window.
It knows exactly what just happened and why it didn't draw well, gradually adjusting its pen usage and end-effector control.
Will Foundation Model Companies Obliterate Embodied Models?
👦🏻 Koji
There's a narrative now that foundation model giants like OpenAI and Anthropic are the only ones who can build the ultimate embodied brain, which would be a dimensional reduction attack on vertical embodied AI startups. What do you think?
🧑🏻💻 Jiawei Wang
An embodied system is like a human — besides the brain responsible for understanding and decision-making, you need a cerebellum controlling muscular reflexes.
GPT-6 can indeed directly manipulate robotic end-effectors, but look closely and the execution frequency is very slow. Single inferences take a long time, and many videos are sped up 16x or 32x. Static environments are fine, but the real physical world is dynamic. Say there's a half-full glass on the table — if the gripper slips during grasping and the cup tilts, the system can barely recover once it starts to spill.
This millisecond-level dynamic adjustment can't rely solely on slow "System 2" thinking from the brain. It requires a cerebellar Action Foundation Model with muscular reflexes.
There's definitely overlap between the brain and large models, but the division of labor is clear: the upper layer uses large models for complex scene understanding and task planning, while the bottom layer must be paired with excellent Action Policy for fast adjustment and agile response.
👦🏻 Koji
Do you think the embodied brain will naturally emerge within language model architectures by adding embodied data, or will entirely new architectures emerge?
🧑🏻💻 Jiawei Wang
The brain side has probably already converged on GPT-like architectures. Their commonsense reasoning and agentic capabilities are extremely powerful. It's not realistic to build a new paradigm from scratch to replace them.
The more realistic path is to push forward along existing lines — converting first-person and third-person interactive videos into trainable task formats to strengthen physical-world understanding. The major foundation model labs are all doing this too.
👦🏻 Koji
Tell me about the Youth Class. How many students per cohort?
🧑🏻💻 Jiawei Wang
Around thirty-something.
👦🏻 Koji
What do these people have in common?
🧑🏻💻 Jiawei Wang
Strong self-drive, quick thinking, and very fast absorption and conversion of new knowledge.

👦🏻 Koji
How does the Youth Class select students?
🧑🏻💻 Jiawei Wang
Students in their first or second year of high school, or even ninth grade, under 15 years old take the gaokao. If their score is within 20 points below the University of Science and Technology of China's local admission line, they advance to the interview.
On interview day, all candidates are gathered together for a university-level math and physics class. Immediately after, they're given an exam — a direct test of focus and the ability to digest and convert knowledge in a short time without prior foundation.
After passing, the dean interviews each candidate one-on-one with open-ended questions, like what if you can't keep up with coursework or want to switch majors.
👦🏻 Koji
Is the interview elimination rate high?
🧑🏻💻 Jiawei Wang
Roughly 100 people enter the second round, and 30 to 40 are selected.
👦🏻 Koji
What's the biggest impact the Youth Class left on you?
🧑🏻💻 Jiawei Wang
The most direct thing was escaping senior year of high school — I never experienced that intensity of endless practice tests.
👦🏻 Koji
Not being poisoned by senior year. Professor Yong Yu of Shanghai Jiao Tong University's ACM Class once said that the first year of university is spent largely detoxing students from the damage of senior year, and only in sophomore year does real cultivation begin.
🧑🏻💻 Jiawei Wang
First year was all foundational general education — math, physics, chemistry, biology, and computer science. You figure out your interests before choosing a major. I picked computer science in my second semester.
USTC has abundant research resources. As an undergrad, I worked with professors on very early AI for Science, using computational simulation to study phase transitions in biological molecular motors.
I was also Youth Class monitor for quite a while, did a lot of student work, and got pretty good at communication. People liked me.
👦🏻 Koji
Very typical ENTJ.
🧑🏻💻 Jiawei Wang
Yeah.
What is real intelligence?
👦🏻 Koji
Growing up surrounded by geniuses, how do you now understand "intelligence"?
🧑🏻💻 Jiawei Wang
My understanding used to be superficial — I thought it meant quick thinking, learning fast. Now I see quick thinkers everywhere; that's not scarce.
I think the core traits of intelligence are two: first, the ability to ask good questions; second, knowing your own weaknesses and being willing to face them head-on.
👦🏻 Koji
What are your weaknesses?
🧑🏻💻 Jiawei Wang
My understanding of hardware isn't deep enough. My father was a computer teacher, and I was intensely interested in software from childhood, but I never got excited about soldering boards or building circuits.
At Simple AI, robotics hardware is an unavoidable required course. I know it's a weak point, so I frequently consult the hardware team and seek their support.
Why leave large models for embodied AI?
👦🏻 Koji
Many people prefer to play to their strengths. You were doing multimodal large model work at DeepSeek and Seed — software-biased work. Switching to embodied AI means confronting massive unfamiliar hardware and engineering challenges. Why make this choice?
🧑🏻💻 Jiawei Wang
As a new grad, I wanted to go to a field where I could define things myself.
Large language models have already converged and are highly industrialized. Young graduates entering big tech are mostly executing predetermined plans — it's hard to explore new directions according to your own ideas. But embodied AI is in a highly unconverged state across the entire industry, with too many critical points waiting to be broken through. When facing something without strong established cognition, I have an instinctive impulse to thoroughly figure it out.
👦🏻 Koji
For most people, unconverged technology means huge risk. But in your eyes it's opportunity?
🧑🏻💻 Jiawei Wang
The bigger the waves, the more expensive the fish. Precisely because there are no standard answers, there's possibility for breakthrough results.
My advisor always encouraged me: young people should go for it, failure doesn't matter, as long as you accumulate enough experience, you can always start the next venture anytime. I have relatively high risk tolerance — risk and opportunity always coexist.
👦🏻 Koji
What's the biggest risk right now?
🧑🏻💻 Jiawei Wang
Time. If the underlying path is wrong, exploring for years only to find you're completely on the wrong track — that kind of blow is the biggest for an individual.
👦🏻 Koji
Among the various explorations in embodied AI, which technical paths are showing emerging risks?
🧑🏻💻 Jiawei Wang
The core issue is that everyone's definition of "intelligence" in embodied AI is vague. A robotic arm following a fixed trajectory to pick up a cup is not intelligence at all.
At Simple AI, we focus on three core capabilities:
First, adaptability, including zero-shot generalization, and when facing completely OOD scenarios, needing only few-shot to rapidly adapt;
Second, steerability. Not just understanding natural language, but also comprehending image and video input — given a human demonstration video, immediately follow it;
Third, native contextual understanding capability. Many embodied models rely only on current single-frame observation, ignoring temporal causality, while physical-world dynamics must emerge through context.
The foundation model is built entirely along these three dimensions.
👦🏻 Koji
Looking back at your PhD years, you happened to catch the rapid development of large models and embodied AI. How do you divide your research stages?
🧑🏻💻 Jiawei Wang
Mainly two stages. The first half was at MSRA in Qiang Huo's group doing document intelligence, also my PhD topic. Professor Huo highly valued practical research value — the OCR, layout analysis, and form understanding algorithms we developed were deployed into Microsoft cloud APIs and internal infrastructure. I got used to thinking about whether academic research could generate actual commercial value upon landing.
The second half, the large model era arrived. I wanted to be a surfer in this wave, and interned at DeepSeek and Seed successively, working on multimodal language models, reinforcement learning, and agents.
DeepSeek: Not Trading Work Intensity for Results
👦🏻 Koji
When you joined DeepSeek, it was after V2's release and before R1 emerged. What was the internal atmosphere like?
🧑🏻💻 Jiawei Wang
The researchers had very pure mindsets — not rushing to think about commercialization, all attention on pushing the boundaries of intelligence.
What impressed me most was that work intensity wasn't actually high, but output was extremely astonishing. This shows that for researchers, long hours of high-intensity work aren't a necessary condition for good results. The team was small and elite, with each person covering a very large scope, able to touch the core training and inference infrastructure code.
👦🏻 Koji
How did you join DeepSeek at the time?
🧑🏻💻 Jiawei Wang
On one hand, a senior from MSRA provided information. On the other hand, they were working on DeepSeek-VL2 at the time, and document processing was a core application direction. HR found me, the background was a good match, and I went.
👦🏻 Koji
The DeepSeek experience was good, so why did you choose to go to Seed later?
🧑🏻💻 Jiawei Wang
At the time, DeepSeek's core resources were倾斜 toward pushing R1 on language models. The multimodal team had relatively limited resources, and continuing would have involved a lot of reusing past experience.
I had already seen agent potential, and Seed internally happened to have a team exploring agents very early, so I decided to intern there.
👦🏻 Koji
Between these two top teams, DeepSeek and Seed, how did the internship experiences differ?
🧑🏻💻 Jiawei Wang
DeepSeek was small and flat, with broader boundaries of responsibility.
Seed was like an extremely massive and precisely operating instrument, with extremely fine professional division of labor. Everyone was anchored at critical points, and it was hard to touch things outside your scope.
This is a double-edged sword: the upside is that engineering teams prepared data infrastructure and training infrastructure very thoroughly — ready to use. The cost is that researchers have relatively weaker control over underlying pipelines.
👦🏻 Koji
What specific event later prompted you to leave pure software large models and leap into embodied AI entrepreneurship?
🧑🏻💻 Jiawei Wang
In late 2025, while interning at Seed, I really wanted to push model self-evolution — deploying models as APIs to serve externally, self-iterating through collecting user feedback.
But under a large company structure, this kind of exploration has extremely high costs. Exposing APIs externally involves multi-department coordination, requires extensive upfront POC work, and exploratory projects can actually get very few resources. Language models have already highly converged; new grads entering find it hard to define things themselves — more about keeping pace with the main force pushing business.
Simple AI happened to be doing full-stack embodied R&D. The founding team gave tremendous freedom, plus the embodied field still has massive unconverged topics waiting for research, so I decided to join.

Why Must Simple AI Do Full-Stack?
👦🏻 Koji
Simple AI does everything from underlying data, models, and hardware to upper-layer agentic OS — the full chain. What's the thinking behind this?
🧑🏻💻 Jiawei Wang
Simple AI positions itself as an end-product company, not a pure data company, not a pure model company.
The embodied industry chain is extremely long. From data, models, deployment to real hardware, missing any intermediate link means no final value can be produced, because what users pay for is the complete product.
Doing full-stack isn't for the sake of doing full-stack — it's that to achieve the goal, you have to do the full chain. Full-stack creates a massive moat: if upstream models suddenly break through in the industry, because we own complete data and real-hardware pipelines, we can quickly follow and iterate into our own products, rapidly getting user feedback from real scenarios.
👦🏻 Koji
Simple AI open-sourced a dataset recently?
🧑🏻💻 Jiawei Wang
Last month we open-sourced the pre-training dataset HiFi-UMI, the cornerstone for training foundation models. Cumulative downloads on Hugging Face and ModelScope have already exceeded 500,000.
👦🏻 Koji
How large was the open-source release?
🧑🏻💻 Jiawei Wang
We open-sourced 2,000 hours of high-quality data. Our internal reserves have reached tens of thousands of hours, accumulated over roughly six months.
At first, we never planned to build such refined data-collection hardware ourselves. We hoped to buy off-the-shelf equipment. But after testing many devices on the market, we found their precision and time synchronization simply couldn't meet pre-training requirements.
👦🏻 Koji
What does your self-developed data-collection hardware look like?
🧑🏻💻 Jiawei Wang
We developed an ultra-high-precision UMI glove. UMI is a universal manipulation interface proposed by Shuran Song's team at Stanford. Data collectors wear the gripper glove to operate, recording trajectories as training corpora.
We solved two critical challenges along the way:
First, completely eliminating base stations. In the past, high-precision trajectory recovery required setting up multiple base stations on the table. Through our self-developed visual methods, we achieved millimeter-level positioning using only onboard cameras.
Second, microsecond-level timestamp synchronization. Many teams working on UMI encounter timestamp desynchronization issues and can only compensate with post-processing algorithms. We guaranteed absolute data precision at the source. Only with this can the model side effectively control variables during experiments.
👦🏻 Koji
How many camera feeds are on the full setup?
🧑🏻💻 Jiawei Wang
Six feeds total: a stereo camera pair on the head; and two pairs of cameras on each wrist pointing up and down. The dual top-and-bottom perspectives prevent visual occlusion during manipulation.
👦🏻 Koji
Why collect so many perspectives? Where's the balance point?
🧑🏻💻 Jiawei Wang
Wrist cameras are essential. When operating with a two-finger gripper for delicate tasks, you must guarantee an unobstructed view—the head-mounted camera alone gets blocked far too easily. The head camera's main purpose is enabling high-precision SLAM spatial positioning to eliminate base stations, and it can be migrated to egocentric data collection in the future. We didn't over-engineer this project.
👦🏻 Koji
How did you initially verify that the collected data had genuine training value?
🧑🏻💻 Jiawei Wang
Very few teams in China have actually gotten UMI working end-to-end. We hit many pitfalls along the way.
After collecting the first batch of data, we fine-tuned on open-source foundation models and ran direct real-hardware evaluations to check task success rates. Once validated, we became confident that with correct collection methodology, we could scale to more tasks.
We also replayed trajectories in simulation to see whether collected trajectories could execute smoothly on the physical robot body. We found over 95% could be reproduced, which solidified our confidence in large-scale collection. So we expanded equipment and leased facilities to scale up.
👦🏻 Koji
What's the current investment ratio between data collection and annotation/cleaning?
🧑🏻💻 Jiawei Wang
Early on, data collection had higher investment. Beyond dedicated facilities, we also did field collection in homes, hotels, and guesthouses to expand the diversity of data distribution.
But recently, investment in annotation and cleaning has overtaken collection. Training directly on rough raw data still works—which proves the data precision is high—but because of mismatches between annotation and actual actions, it seriously impacts the model's learning efficiency. So our current focus is building standardized, automated data-cleaning pipelines.
👦🏻 Koji
Does Astra make automatic annotation possible?
🧑🏻💻 Jiawei Wang
We tried using GPT-6 Astra for a batch of cleaning and annotation, and found it actually outperforms human annotators. It can run 24/7 at roughly the same cost as human labor.
The only constraints are API quotas and per-batch efficiency. So we're using high-quality data annotated by something like Astra to train our own annotation small models, iterating continuously in small batches. The results are already quite good.
👦🏻 Koji
What was your thinking behind open-sourcing 2,000 hours of heavily invested data?
🧑🏻💻 Jiawei Wang
People with research backgrounds all have an open-source ethos—they want their work to be seen and used by the community, advancing the field. Two thousand hours is already quite practical for university labs.
On the other hand, we also hope to collect feedback from the community to continuously iterate our data-collection hardware.
👦🏻 Koji
What feedback have you received since open-sourcing?
🧑🏻💻 Jiawei Wang
Many domestic and international companies have emailed wanting to buy data.
On the technical side, two main points: First, the QR code markers used for hand positioning are too large, which interferes with visual algorithms—people want them smaller. Second, the gloves get noticeably warm during extended real-hardware operation, which relates to camera module suppliers. This feedback has been very practical.
Can Robots Really Zero-Shot?
👦🏻 Koji
What progress can you share early about the model you'll release by year-end?
🧑🏻💻 Jiawei Wang
After pre-training, the model has already demonstrated strong Zero-Shot generalization capabilities.
Previously, very few models in the industry could do direct Zero-Shot. Most relied on real-hardware post-training to benchmark-tune scores. We do Zero-Shot validation to see the model's underlying true capabilities. After pre-training on thousands of tasks, without any post-training fine-tuning, it can directly generalize to complete some entirely new tasks.
👦🏻 Koji
Similar to GEN-1.5's ability to learn from watching a 3-second video?
🧑🏻💻 Jiawei Wang
GEN-1.5 is One-Shot—it needs a reference video. We have certain Zero-Shot capabilities. Our internal discussions attribute this core ability to the extremely high quality of our underlying pre-training data.
👦🏻 Koji
Are you building a world model?
🧑🏻💻 Jiawei Wang
It depends on how you define world model. We don't deliberately chase concepts, but rather work backward from core capabilities to architecture: whether World Action Model or VLA, as long as there's a suitable capability point, we'll absorb it into our model design.
We're not a follow-the-leader shop. We design models most suitable for physical intelligence from first principles, while fully leveraging existing pre-trained knowledge.
👦🏻 Koji
You just emphasized that one of the model's three foundational capabilities is Context. Why is Context so important?
🧑🏻💻 Jiawei Wang
Anyone working on language models and Agents knows Context is core to decision-making. But in embodied AI, many people have assumed that decisions can be made from single-frame observation alone, ignoring the physical causality embedded in Context.
Embodied Context has at least three levels:
Short-term Context, containing high-frequency causality. For example, the first time you grip a coffee cup, it slips because you didn't clamp tight enough, then you apply more force to restabilize it. The visual slip and re-clamping closed-loop that happens within that one second is rich causal Context. The model naturally learns what went wrong just now and how to correct it next time.
Medium-term Context, more like working memory. There are several identical cups on the table. After trial-and-error grabbing the first one, by the second and third, the model already knows the cup width and can directly transfer without repeated trial-and-error.
Long-term Context, memory for extended tasks. When doing housework, you put something in a drawer half an hour ago. When a human gives an instruction, you can immediately locate it without needing reminders every time.
👦🏻 Koji
What designs have you made for Context at the data and model levels?

🧑🏻💻 Jiawei Wang
At the data level, you need to construct samples that stimulate Context utilization; at the model level, you need to efficiently model Context.
Long Context brings massive token counts, causing inference latency to rise. Robots are extremely sensitive to execution latency. So we've made special designs in temporal modeling and attention mechanisms to address inference latency and training efficiency.
👦🏻 Koji
The HiFi-UMI dataset doesn't include tactile data?
🧑🏻💻 Jiawei Wang
Some directions now try predicting visuo-tactile sensing directly from video, but the precision isn't sufficient for real-hardware control.
There are two engineering considerations for why the open-source data doesn't include tactile:
First, for a two-finger gripper, combined with the blind-spot-free design of top-and-bottom dual cameras, vision can already judge whether the grip is tight in the vast majority of cases—unlike dexterous hands which have severe palm-internal occlusion.
Second, tactile sensors haven't converged in industry yet; resolution and placement vary widely. If you spend heavily to collect tens of thousands of hours of data, then hardware iterates months later, previous data becomes hard to reuse. We've been experimenting with tactile, but haven't rashly scaled it up.
Has the Embodied Scaling Law Emerged?
👦🏻 Koji
Pretty interesting. I recently recorded another podcast with a guest who's a quantum dot scientist—his PhD advisor won the Nobel Prize in Chemistry for quantum dot technology. What they do is emit quantum dots to sense touch.
Back to models—you mentioned seeing some Zero-Shot capabilities. Do you think this is a signal of Scaling Law?
🧑🏻💻 Jiawei Wang
Scaling Law is a very serious matter. We have indeed observed scaling trends: as data diversity, data volume, and compute expand, performance improves significantly.
But a Law in the strict sense requires rigorous experimental validation, quantitatively determining the proportional relationships between data volume, model scale, and compute to make training predictable.
Many teams in the industry like to claim they've found Scaling Law. We more conservatively call it an observed Power Law.
👦🏻 Koji
How does this trend manifest in specific tasks?
🧑🏻💻 Jiawei Wang
After pre-training on one to two thousand tasks, we originally assumed only head tasks with large sample sizes would generalize. Later we found that low-frequency data representing only 0.1% to 0.5% of the total could also perform well.
Additionally, cross-task compositional generalization emerged. For example, the training set only had "put the remote in the box" and "put snacks in the trash can." The model had never seen "put snacks in the plate." But given the new instruction "put snacks in the blue plate," it could compose and complete this directly without any post-training.
Once data diversity reaches a certain point, model capabilities show jumps beyond imagination.
👦🏻 Koji
What architecture is Simple AI's Agentic OS?
🧑🏻💻 Jiawei Wang
Embodiment is necessarily an Agent System, not a single-point model. It needs to interact with humans and the physical environment. The core is building a good Harness for embodied scenarios.
It includes speech recognition, task decomposition, tool calling, and so on, divided into two layers: System 2 is the brain, handling semantic understanding and long-horizon planning; System 1 is the cerebellum, handling high-frequency precise execution. We've done extensive Harness design around System 2—for example, at WRC, users can use natural voice multi-turn dialogue to have it find clothes in a room and put them in the washing machine.
👦🏻 Koji
What model is behind System 2?
🧑🏻💻 Jiawei Wang
We were using ByteDance Seed's model before, since I'd worked on that and knew it well.
👦🏻 Koji
Will you switch to a self-developed model later?
🧑🏻💻 Jiawei Wang
We're trying it out. Brain models consume enormous compute, and for a startup, compute is extremely precious. Right now we're prioritizing pre-training for the Action Foundation Model.
General-purpose large language models are iterating incredibly fast — the leap from GPT-5 to GPT-6 is huge. Our current strategy is to make the cerebellum a highly Steerable, tightly controllable interface. With good Harness design, we can leverage the cutting-edge brain capabilities from the industry and feed that back into commercializing the whole system.
👦🏻 Koji
Will System 2 ultimately just be an extension of frontier lab language models?
🧑🏻💻 Jiawei Wang
Beyond the base model's capabilities themselves, what matters most in embodied scenarios is the Harness.
The brain receives high-resolution video streams in real time. If context management isn't handled well, the context window explodes immediately and latency becomes unbearable. The interaction protocol between brain and cerebellum also needs deep co-design — only when the cerebellum is made extremely controllable can the brain schedule it naturally.
No Benchmark — how do you prove the model?
👦🏻 Koji
Embodied model evaluation has long lacked good Benchmarks. What's your take on the current state?
🧑🏻💻 Jiawei Wang
Past Benchmarks really haven't been great enough. Pure simulation benchmarks like LIBERO can be gamed to extremely high scores with pure post-training fine-tuning, so they don't mean much.
Later we got things like RoboDojo that combine simulation with real hardware, but simulation has a Sim-to-Real Gap, and real hardware is hard to reproduce across different robot platforms — third parties can't easily run someone else's model on their own hardware.
Embodied evaluation tests the stability and success rate of the entire pipeline on long-horizon tasks. The most solid approach right now is building rigorous real-hardware test sets yourself.
👦🏻 Koji
Without fair industry standards, how do you plan to prove your strength when you release your model at year-end?
🧑🏻💻 Jiawei Wang
The only way is to stay honest.
Build rigorous real-hardware Benchmarks, from simple to complex, testing generalization across scenarios and objects under perturbations, and publish real success rates transparently. Report exactly how much fine-tuning data was used. If you label something Zero-Shot, don't sneak in a single human demonstration.
Different teams have different hardware platforms and data, making direct apples-to-apples comparison difficult. The most important thing is honesty. If there's an open-source plan, we'd hope to release a model that third parties can take and fine-tune cheaply on their own hardware to get good results.
👦🏻 Koji
Will System 1 and System 2 eventually converge into a unified end-to-end large model, or stay layered long-term?
🧑🏻💻 Jiawei Wang
The lower layer swallowing the upper layer seems unlikely, but the upper layer swallowing System 1 is genuinely possible — though it's hard to judge exactly when that inflection point might come.
Even if you can see the endgame, you can't skip the steps you have to take along the way. Big tech can afford to try building a multi-trillion-parameter end-to-end unified model. For a startup, the most pragmatic strategy is to solidify Action capabilities first, make the cerebellum extremely Steerable, pair it with upper-layer large model planning and Harness, and get to commercial deployment in real scenarios quickly.

👦🏻 Koji
How big is Simple AI's team now?
🧑🏻💻 Jiawei Wang
Over 70 people.
👦🏻 Koji
At DeepSeek you noticed people left work early — excellent results don't scale linearly with hours worked. Does Simple AI have a lot of overtime?
🧑🏻💻 Jiawei Wang
No hard overtime requirements. It's all self-driven. Normal hours are 9:30 to 6:30. Leave early if you want or need to — I don't work late every night myself. Staying mentally sharp matters.
With Agent support, everyone needs to manage their own Agents more, clearly defining task objectives to push forward. The company puts no limits on AI tool budgets.
👦🏻 Koji
What Agent tools do you mainly use internally?
🧑🏻💻 Jiawei Wang
Mainly Codex.
👦🏻 Koji
Any special meeting culture internally?
🧑🏻💻 Jiawei Wang
Emphasis on planning ahead. The CEO regularly pulls together core tech leads for top-down reasoning: What is the intelligence in Embodied AI? How do we prepare data? How do we match compute? The team exchanges context thoroughly, sharing information flatly so it doesn't get bottlenecked with just a few people.
👦🏻 Koji
What's the biggest takeaway from GPT-6 Astra?
🧑🏻💻 Jiawei Wang
General intelligence really could spill over into the physical world. The boundary between System 1 and System 2 will get increasingly blurry.
Practitioners coming from the large language model perspective believe it's entirely possible to make this genuinely usable. Though inference is slow now, as long as you incorporate physical interaction Action data, that gap can be closed.
👦🏻 Koji
Domestically, who are you most bullish on for building Astra-like physical interaction capabilities?
🧑🏻💻 Jiawei Wang
Two teams: one is ByteDance Seed's multimodal team — they've long had leading multimodal foundations; the other is Alibaba's Qwen team, which is doing very deep work internally on the embodied direction.
👦🏻 Koji
Globally, what embodied releases from the past six months have you found most inspiring?
🧑🏻💻 Jiawei Wang
Most respect definitely goes to Generalist.
I think their approach is exactly right: GEN-0 validated Scaling with 200,000 hours of data; GEN-1 reached 500,000 hours and generalized across 9,000 gripper configurations; GEN-1.5 showed stunning In-Context Learning.
Also Professor Levine's work on π0.7, which puts Steerability at the core and proves that only a highly controllable cerebellum can cooperate with the brain to complete complex tasks.
👦🏻 Koji
Any tech that's blown up on social media but you think is overrated?
🧑🏻💻 Jiawei Wang
Some teams' In-Context Learning based on Paired data.
Generalist's In-Context Learning was stunning because it emerged naturally from training on 30-second-long Context data — give it a very short third-person human demonstration video and it can follow directly.
But some teams'跟风 Demo releases were hard-trained on specially constructed paired data. That's just an engineering compromise, lacking real substance. Genuine In-Context Learning must be built on deep understanding of Context and strong model Steerability. Without these two, Demos brute-forced out don't mean much.
👦🏻 Koji
Embodied models are hard for others to reproduce — will this lead to third-party neutral Benchmarks emerging?
🧑🏻💻 Jiawei Wang
Hard. Some institutions have tried building real-hardware Benchmarks, but they ended up both referee and player, and it fizzled out.
Embodied model evaluation is difficult to guarantee as absolutely fair. In the end, it comes down to product — whether users buy in.
👦🏻 Koji
You've been deep in embodied for almost a year, and much longer in large models before that. What's fundamentally similar and different between these two fields?
🧑🏻💻 Jiawei Wang
Large language models are overall more open and transparent. Research is relatively reproducible, there are APIs to test, and DeepSeek fully open-sources its technical breakthroughs — barriers are relatively transparent.
Embodied currently has more hype than reality, too much noise. In the first half of the year everyone was touting World Action Model, but very few have actually landed and run business operations with it. There's no clear conclusion on how much better it really is compared to traditional VLA — easy to get caught up in the froth.
👦🏻 Koji
In all this industry noise, how do you filter for real signal?
🧑🏻💻 Jiawei Wang
Two things: first, whether the claims match the Demo details presented; second, stripped of packaging, whether it aligns with first principles.
Precisely because embodied hasn't fully converged, we can absolutely do the right thing starting from first principles. For example, self-developing high-precision data collection hardware — it's heavy, but controlling data quality at the source actually let us leapfrog.
👦🏻 Koji
What is Simple AI's first principle?
🧑🏻💻 Jiawei Wang
Keep asking: what exactly is the intelligence in Embodied AI?
A robotic arm following a fixed trajectory to grab a cup isn't intelligence. The ability to freely perceive, understand, and achieve goals despite lighting perturbations, unseen objects, physical slippage, and novel instructions — that is Embodied AI.
👦🏻 Koji
Any hotly hyped approach that you actively abandoned after deep reasoning?
🧑🏻💻 Jiawei Wang
Not quite abandoned, but we slowed down on large-scale deployment of Egocentric first-person human video data.
Ego data is extremely diverse, but doing the first-principles math: raw data costs only 30 yuan per hour, but the subsequent annotation, cleaning, and model fine-tuning compute costs to map human motion to robotic arms are extraordinarily expensive. Whether the generalization payoff slope is worth it hasn't been rigorously validated.
UMI data has a self-developed hardware barrier, but self-collection costs are controllable, over 90% of data can be directly used for high-quality training, and you can clearly validate parameter-data scaling laws with lightweight compute.
At a resource-constrained stage, blindly introducing big variables you haven't accounted for is very dangerous.
👦🏻 Koji
This suddenly reminds me — at the start when I asked how you understand "smart," you said a core part is "asking good questions." What I worry most about doing podcasts is my Prompt not being good enough, not drawing out valuable information from guests.
If any listeners feel there were places in our conversation where a better question could have been asked, feel free to leave a comment — I'll pass it to Jiawei to continue the exchange.
Now I'll hand the questioning over to you — anything you're curious to ask me?
🧑🏻💻 Jiawei Wang
There's one interesting question I've always been curious about, probably not related to embodied: What are the similarities and differences between doing podcasts and doing investment? I've toyed with the idea of podcasting before — being able to have deep conversations with people across industries seems fascinating.
👦🏻 Koji
Absolutely go for it — even recording on your phone in a café, the audio quality is good enough.
The biggest similarity between podcasting and investing is this: you get two hours to ask someone questions and try to really understand them. For instance, I always ask, "What's the hardest decision you've made at a crossroads in your life?"
When I'm podcasting, I love hearing stories; when I'm investing, I'm equally desperate to see, through choices made in extreme difficulty, whether someone has that founder personality.
The difference lies in the state of mind: when podcasting, I'm very nice — I avoid questions that might make the guest uncomfortable, wanting them to stay in a comfortable emotional space where they can express themselves freely.
When investing, the questions have to be sharp. Even if I know it'll make the other person uncomfortable, I have to keep pushing, drive them into an extreme state, because only that kind of reaction yields more honest answers.
Back to you: from the youth talent class to a top-tier institution to chief scientist, your path looks incredibly smooth — the quintessential "other people's kid." What's the hardest crossroads decision you've ever faced?
🧑🏻💻 Jiawei Wang
The hardest decision was whether to stay in the comfort zone of language models or pivot to embodied AI.
Giving up the generous salary and GPU resources at a major tech company to join a startup and build infrastructure, design models, and assemble a team from scratch — stepping out of my comfort zone to face complete uncertainty. I was extremely torn at the time.

👦🏻 Koji
What are your former classmates from the youth talent class doing now?
🧑🏻💻 Jiawei Wang
Quite a few went to the United States for PhDs in foundational disciplines like math and physics. But even among those who studied foundational sciences, many have pivoted to AI — either working at the intersection of AI and their field, or in AI-driven quantitative finance, AI healthcare, or AI drug discovery. One classmate even went to OpenAI to do theoretical research.
👦🏻 Koji
Has anyone made a non-mainstream, non-consensus choice — completely avoiding AI, not going into finance, and doing something else entirely?
🧑🏻💻 Jiawei Wang
Almost no one. Under the massive trend of artificial intelligence, everyone is embracing it.
Will robots be able to do dishes in three years?
👦🏻 Koji
AI has absorbed all the top talent. We often say people overestimate change in two years while underestimating change in ten. Three years is an interesting timeframe — looking back from three years in the future, what event would make you feel that entering the embodied AI industry was completely the right call?
🧑🏻💻 Jiawei Wang
My doctoral advisor once mentioned a pain point: even with a dishwasher at home, you still have to manually scrape leftover food into the trash and load the plates one by one, which takes about twenty minutes every day.
If three years from now, our self-developed robot can reliably handle tasks like clearing scraps and loading the dishwasher in a home setting, I'd be incredibly excited.
👦🏻 Koji
What's the probability of achieving that goal within three years?
🧑🏻💻 Jiawei Wang
Pretty high — I'd say 80%.
👦🏻 Koji
Really looking forward to the robot from Simple AI that'll do the dishes for my family in three years. Thanks, Jiawei!
🧑🏻💻 Jiawei Wang
Thanks, Koji!
Join the membership group
