World Models and Intelligent Infrastructure: Deep Dive into the New Wave of Embodied Artificial Intelligence, Different Answers Emerge | Yunqi Capital × WRC
First-Hand Judgments from Frontier Explorers

How can robots truly understand the world? And what kind of foundation should underpin that intelligence?
At the recently concluded 2026 World Robot Conference, a dense lineup of new products, technologies, and applications showcased the present state of Embodied Artificial Intelligence. Meanwhile, a series of cutting-edge presentations and spirited debates continued to open up new possibilities for its future.
On the afternoon of August 21, the special event "The New Wave of Embodied AI: World Models and Intelligent Infrastructure," co-hosted by Yunqi Capital and SEE Fund and supported by Mars Accelerator and the Tsinghua University Alumni Association Electronic Engineering Department Branch, was successfully held. This marked Yunqi's fifth consecutive year at the World Robot Conference, engaging with industry practitioners on the latest shifts in embodied AI.
The event focused on two hot topics: world models and intelligent infrastructure. Guests gathered from frontier research teams at universities, domestic leading large model and cloud service platforms, and startups working on world models, robot bodies and applications, and data and computing infrastructure. The event also drew over 200 attendees from across the embodied AI industry, research, and investment communities. Over three hours, discussions ranged from technical approaches to data efficiency, compute infrastructure, and real-world deployment — with speakers and audience alike exploring what it will take for embodied AI to reach its next stage.

Before Consensus Forms, How Do You Judge Direction?
New hotspots and new teams keep emerging, yet technical approaches remain far from unified. How to judge direction? This was a shared concern of co-hosts Yunqi Capital and SEE Fund. Both are among the earliest domestic investors to commit to the embodied AI wave, and in recent years have built out systematic positions around robot bodies and applications, world models, and infrastructure for data and compute.
"As early-stage investors, we're quite accustomed to chasing directions and stages where the answers haven't yet formed and technical routes haven't unified. Many of the most important technological shifts are actually most valuable before industry consensus takes hold."
Yi Han, Executive Director at Yunqi Capital, explained in his opening remarks that since its founding in 2014, Yunqi has remained focused on early-stage tech investing, accompanying over 200 tech startups through their growth over the past decade-plus. Long before embodied AI became a hot topic in venture capital markets, Yunqi had already begun investing in related directions and gradually built out a systematic layout.
He noted that the robot industry's focus is shifting from how to perceive the world and gain stronger locomotion capabilities, to how to truly understand the physical world. This requires robots not merely to recognize what they see, but to comprehend spatial relationships, causal relationships, and physical laws — and to make predictions and adjustments. It is from this requirement that world models have drawn broad attention.
But whether VLA or world models, "we're far from having unified answers." Yi Han pointed out that beyond models themselves, data collection, compute support, and how to combine simulation with real-world data are equally critical to embodied AI's next steps.

Yi Han, Executive Director, Yunqi Capital
With technical routes yet to converge, startups and capital have already poured in rapidly. How to view this surge?
Ma Lin, Managing Partner of SEE Fund, said in his opening remarks that SEE Fund began investing in the embodied AI track in 2023 and has already established a产业链布局. While the embodied AI market has seen active fundraising this year, what matters more than short-term valuations is whether companies are on the right path and still doing the right things.
From recent World Robot Conferences, Spring Festival Galas, and robotics competitions, robot capabilities have continued to improve. "The transformation that embodied AI and robotics will bring to industry — and to the world — is no longer distant." But he candidly added that while general embodied AI is the goal, it "can't be achieved overnight." Industry规律 must still be respected, and progress must be steady.
Embodied AI will provide incremental value for humanity — doing what humans cannot, and reaching where humans have not. This is the direction that deserves more attention from the entire industry.

Ma Lin, Managing Partner, SEE Fund
From the Front Lines:
How AI Understands the Physical World
With technical paths yet to converge, explorations in different directions have already gone deep into embodied model training and real-world scenarios. In keynote presentations, three guests drew on frontline practice to share how large model methods are entering embodied scenarios, how physical intuition can be learned efficiently, and how world models can achieve physical trustworthiness.

How Do Large Model Methods
Enter Embodied Scenarios?
Drawing on large model platforms and industrial practice, Yang Liwei, Vice President of Volcano Engine, introduced the Doubao model family's explorations in real-time multimodal interaction, connecting perception and execution, and post-training and reinforcement learning, using cases such as museum guides.
Methods already validated in the large model domain are beginning to enter embodied scenarios. Yang Liwei took In-Context Learning as an example: large language models can adapt to new tasks without updating parameters, simply by reading a few examples. Embodied intelligence can draw on this approach, but needs to enhance the richness of context modalities, temporal length, and window capacity.
Data efficiency is equally critical. Current UMI data has already reached millions or even tens of millions of hours, without necessarily bringing corresponding capability improvements. Yang Liwei believes that improving data quality still requires solving problems including cross-modal representation unification, real-world scenario fidelity, and extracting prior information beneficial to downstream actions.

Yang Liwei, Vice President, Volcano Engine

Beyond "Accumulating Assets,"
Raising Scaling Efficiency
Yonglu Li, Associate Professor at the School of Artificial Intelligence, Shanghai Jiao Tong University, presented "World Model and Physical Intuition Learning," introducing his team's research on training models to understand geometric relationships, physical changes, and long-horizon tasks using real-world data, internet videos, games, and simulation.
In his view, imitation learning and reinforcement learning have formed relatively mature paradigms, with increasing technical convergence across the industry — "this shows the industry has reached consensus." But effective methods still require continuous "asset accumulation." After data enters the millions or tens of millions of hours, storage, retrieval, and distributed training also extend competition into infrastructure capabilities.
However, "hours" may not be the best unit for measuring data value; what matters is whether data improves model capabilities. Yonglu Li introduced that his team is also experimenting with using low-cost data from games and simulation to train models' "physical intuition," and measuring generalization through cross-task performance and the amount of data needed to adapt to new tasks.

Yonglu Li, Associate Professor, School of Artificial Intelligence, Shanghai Jiao Tong University

How Do World Models Clear
The "Physical Trustworthiness" Hurdle?
A world model that "looks real" — does that mean it conforms to physical laws and can truly serve robot action? Dingkang Yang, Co-founder and CTO of Feijiekesi, a Yunqi Capital early-stage portfolio company, shared his team's explorations in "Creating, Understanding, and Simulating the Physical World."
He analyzed several mainstream approaches to world models: video generation routes excel at visual presentation but may not conform to real dynamics; physical knowledge in latent variables is difficult to verify; 3D/4D models can reconstruct and roam scenes, but "being roamable doesn't mean being interactive or simulatable." Dingkang Yang believes that for world models to enter the physical world, they still need to add simulation and inference capabilities based on dynamics and physical properties.
When truly facing robots, a physical world model cannot merely predict "what the next frame looks like" — it must also judge what results an action will bring. Along this line of thinking, Feijiekesi released Fysiverse in July this year, using an "explicit physics simulator + generative renderer" to establish computable physical states and simulate future changes that conform to physical laws.

Dingkang Yang, Co-founder and CTO, Feijiekesi
Report Release:
Who Can Turn the Real World Into Verifiable Experience?
How models learn and understand the physical world cannot be separated from data and infrastructure. At this special event, Mars Accelerator, Yunqi Capital, and SEE Fund jointly released the 2026 Embodied AI Data Industry Chain and New Infrastructure Research Report, providing a stage-by-stage梳理 of industry status from perspectives including data infrastructure, technical evolution, and investment and financing trends.
At the event, Fen Li, Co-founder of Mars Accelerator, shared some insights from the report in a presentation titled "Data Flywheel: From Handicraft Workshop to Experience Factory."
Fen Li noted that getting robots onto real production lines requires technical teams to be stationed long-term and debug repeatedly — early deployment is not only "labor-intensive" but "PhD-intensive." Even a simple action requires task design, synchronized collection, human-robot mapping, and data governance. Thus, "one action does not equal one valid experience." What matters is not how much data was collected, but how much robot capability was generated.
Only by bringing failure samples and edge cases from real deployment back into data production, and letting the process iterate continuously, can the data loop become a flywheel. World models may enable experience production to shift from "passive recording" to "active design," but this still requires support from data, compute, simulation, and engineering systems. "Whoever can continuously turn the real world into verifiable experience will master the new infrastructure."

Fen Li, Co-founder, Mars Accelerator
World Models:
How Will Paradigms Evolve?
From model training to data flywheels, embodied AI is becoming a competition of systematic capabilities. The consensus and disagreements within this were also集中呈现 in two roundtable discussions.
The first roundtable took the title "Different Routes to World Models: How Will Robotics Technical Paradigms Evolve?" Yu Sang, Executive Director at Yunqi Capital, was joined by the founders of three embodied AI companies Yunqi invested in during 2026 — Huazhe Xu of Poke Robotics, Sam Li of Quantum Power, and Nianqiu Liu of Feimoyi Technology — for a discussion on technical evolution and industry落地.

What makes a world model truly "useful"? Huazhe Xu believes the key lies not only in predicting the next frame, but in whether it can help robots learn representations, evaluate policies, and complete general manipulation. Sam Li proposed that predicting the world ultimately serves changing the world — models must also combine business semantics to judge whether tasks are completed. Nianqiu Liu argued that robots need to handle not just environmental prediction, but also more complex capabilities including cognition, reasoning, and human-robot interaction — narrow world models alone are insufficient to support autonomous robot action in open environments.
How can embodied AI accelerate Scaling? A consensus in the discussion: Scaling cannot be simply equated with increasing parameters or piling up data hours. Compared to repetitive data, diversity in objects, environments, interaction methods, and failure samples matters more — "success may be千篇一律, but failures each have their own characteristics."
From demo to scaled deployment, the model is only one link. Huazhe Xu focuses on generalization capabilities in unstructured scenarios such as homes; Sam Li believes end-effector reliability, anomaly detection, and autonomous recovery are equally critical system capabilities; Nianqiu Liu sees mobility as the foundation for robots to enter open, dynamic scenarios — only with the ability to act in real environments can robots move toward broader task applications.
The roundtable also touched on embodied talent and the industry landscape over the next three years. Whether investors or entrepreneurs, what they look forward to is not just better demos, but robots truly entering homes, logistics, and more scenarios to create usable value.
Intelligent Infrastructure:
Bottlenecks Aren't Just Compute
The second roundtable took the title "Intelligent Infrastructure: What's Next for Data, Compute, and Infra Innovation?" Moderated by Jingdong Cai, Senior Investment Manager at SEE Fund, the discussion included Yonglu Li, Associate Professor at the School of Artificial Intelligence, Shanghai Jiao Tong University; Xinyu Zeng, CTO of Modou Link; Liruo Zhong, President of X-Era Lab (Tuoyuan Wisdom); and Quanlu Zhang, VP of Technology at Infinigence AI.

What does the industry most lack right now? All four guests pointed to data. Yonglu Li is concerned with the efficiency of acquiring valid data: of collected data, perhaps only 5–10% actually enters training. Xinyu Zeng used "crude oil" versus "refined oil" to describe the gap — large amounts of raw data still need to pass through standardized data pipelines before models can actually use them.
Liruo Zhong noted that her team has accumulated tens of millions of hours of data, yet only one hundred to two hundred thousand hours actually enter the training pipeline. Quanlu Zhang proposed that beyond pursuing higher-quality data, new training methods should also reduce models' requirements for data quality, and let data continuously flow back between collection, training, and deployment.
Embodied infrastructure cannot simply copy solutions from the large model era. Yonglu Li pointed out that robots lack ready-made datasets and unified benchmarks, requiring teams to build their own pipelines connecting data, training, evaluation, and feedback; Xinyu Zeng believes that multimodal data including vision, point clouds, touch, and action also creates new training bottlenecks in storage, alignment, decoding, and loading.
How models enter robots is another key question. Liruo Zhong judges that models will ultimately need to "grow on the robot body itself," making edge deployment and operating systems connecting different bodies more important. Quanlu Zhang focuses on cloud-edge-device collaboration: robots are more sensitive to latency, and real machines must also enter training and iteration — infrastructure must simultaneously connect compute and devices at different levels.
The discussion also extended to Scaling, standardization, and deployment. The exchange of views made the轮廓 of intelligent infrastructure more concrete: it is not any single-point tool, but a systematic engineering effort running through data collection, model training, real-machine deployment, and feedback loops.

How can robots truly understand the world? And what kind of foundation should underpin that intelligence? Three-plus hours of presentations and discussion did not produce a standard answer, but showed us that embodied AI's next steps hold more than one possibility.
Technical routes are still evolving, industry consensus has not yet formed, and true answers will need to be repeatedly verified in real tasks and scenarios. The discussion has paused for now, but exploration continues. Yunqi will also walk alongside entrepreneurs, researchers, and industry partners, accompanying embodied AI as it steps by step enters the real world.
For the full discussion, watch the replay on the "Yunqi Capital" WeChat Channels live stream





