VLA and world models have been hotly debated — did this four-hour tech salon reach any conclusions?
In a rapidly evolving field, sustained observation.

Where are robotics and embodied intelligence headed? Right now, VLA and world models have emerged as two prominent technical approaches drawing significant attention. Are they diverging paths, or will they converge? Progress in the field continues to accelerate, yet no definitive answer has emerged on which model route to take, and as technology moves from labs into the real world and industry, every question and possibility remains on the table.
Last Saturday, BlueRun Ventures' "Booming" series partnered with AgiBot's A-Plan program to convene technical leads from various research institutions and tech enthusiasts in Beijing for an in-person technical salon titled "VLA and World Models: Either/Or, or a Twin Era?"
Cao Wei, partner at BlueRun Ventures, noted during the salon: "Embodied intelligence and robotics are both ultra-long-cycle, ultra-large-scale mega-opportunities." This means we need to stay open-minded and embrace more innovative possibilities. Over four hours, the discussion kept spiraling into deeper questions — from VLA and world models to human-centric data, future factories and homes, and industrial value.
BlueRun Ventures has long focused on robotics and embodied intelligence, building deep understanding and making early-stage investments. The firm believes that at this juncture, what matters is continuously understanding the direction of technical evolution, and through open exchange, gradually identifying the key variables that will shape industrial development.
Below is a detailed recap of the salon. If you're also following embodied intelligence, VLA, world models, and related fields, we welcome you to continue this discussion with us.

"While some early-mover teams have already emerged in VLA and world models today, the industry hasn't yet formed genuine consensus," Cao Wei of BlueRun Ventures said at the salon's opening.
In embodied intelligence, VLA (Vision-Language-Action models) and world models are two technical approaches currently generating intense debate. VLA focuses on "task completion" — directly mapping vision, language, and other inputs into physical actions, with fast response and low latency. World models, by contrast, emphasize a robot's understanding, prediction, and simulation of the real world, hoping robots can adjust their action strategies through reasoning. The central question: Will these two model types develop in different directions, or will they complement and merge?
In Cao Wei's view, embodied intelligence remains at a very early stage of its technical stack. Questions about continuous learning capabilities, real-world data compression, and more have yet to find unified answers. At the same time, technical development in embodied intelligence is in a sustained period of accelerating growth, so there's still substantial content worth exploring and room for innovation.
The salon began with research topics from technical practitioners, then extended during Q&A and open discussion to real-world technology deployment, generalization capabilities, and future directions for embodied intelligence.


Zhang Jianke from Jianyu Chen's research group at Tsinghua University's Institute for Interdisciplinary Information Sciences started from VLA model architecture, discussing the problem of visual-language understanding degradation that occurs when migrating from VLM to VLA. The team's proposed solution, UIM (Unified Interaction Model), attempts to decouple different information types into "what" versus "where and how" capabilities for learning — preserving visual-language understanding while further enhancing robotic action capabilities.
Zhenguo Sun, co-founder of StarGate Intelligence and head of the Embodied Interactive World Model Research Center at BAAI, shared his perspective on the relationship between VLA and world models, arguing they are not simply substitutive but symbiotic. VLA excels at learning concrete actions from successful trajectories, while world models can introduce richer training signals for robots by predicting environmental changes and incorporating more interaction information. Based on this, his team proposed the ω-EVA framework, aiming to jointly model prediction and action. He also noted that the world model technical route has yet to converge, with future exploration needed on model paradigms, multimodal information fusion, and robotic capability enhancement.
Kevin, co-founder of MoKe Robotics, shared his team's series of works around world models, focusing on world model reasoning frameworks — such as their proposed HDR (Hierarchical Diffusion Forcing) — hoping to improve world model stability in long-horizon reasoning, with validation across multiple benchmarks. Additionally, he shared the team's thinking on data paradigms, noting the clear gap between real robot execution data and internet video, and therefore emphasizing first-person perspective (Ego-centric) real interaction data.
Professor Yao Feng, who will soon join Tsinghua University's School of Artificial Intelligence, shared research and reflections on human-centric AI. She believes that as embodied intelligence develops, AI foundation models are gradually shifting from understanding the semantic world to understanding the physical world. Thus, rather than relying solely on text and images from the internet, future models will need to understand human poses, movements, and human-environment interactions. Her research aims to build foundation models capable of perceiving, understanding, and naturally interacting with humans.
Zhao Hao, assistant professor and doctoral advisor at Tsinghua University's Institute for AI Industry Research (AIR), and BAAI Scholar, discussed perspectives on world models and multimodal foundation models. He believes that as model capabilities improve, post-training paradigms like reinforcement learning are becoming important paths for world model continuous evolution, and multimodal foundation models with stronger understanding, reasoning, and interaction capabilities could become important foundations for next-generation embodied intelligence.

Different research teams choose different model representations, which relates to how robots understand the real world. Several technical practitioners offered their perspectives.
Kevin: His team focuses more on extracting semantically informative representations from video, hoping to learn more compact feature representations through pre-trained encoders, and mentioned recent CVPR work on Delta Tokens, aiming to better represent motion information using differences between two frames.
Yao Feng: What a model needs to model depends on the ultimate application goal: industrial robots need more precise understanding of physical properties, while robots for home and living scenarios need to model more object appearances and human-environment interaction information.
Zhao Hao: In large-scale deep learning, discrete representations still hold advantages. An important future research direction is further integrating mechanical information into 3D representations, while continuing to explore motion representation approaches.
Data is also key to supporting models. Kevin noted that in his team's data selection, real robot teleoperation data collection is costly and limited in quality, while mixing Ego-centric data with real robot data currently achieves good results.
Different application scenarios also affect data strategy: if targeting living scenarios, data sources can be richer and more diverse; if focusing on specific tasks and embodiments, data diversity needs to be controlled from the embodiment perspective. More data isn't always better — it needs to serve model generalization and real task closed loops.

Research teams have different model architectures. Will architectures evolve through application scenario migration, or will a unified architecture emerge? On this widely discussed question, the technical practitioners present offered differing views.
Kevin: Drawing on the development history of image generation models, excellent model architectures often emerge through continuous trial and error, and ultimately what matters more is having efficient data infrastructure and continuous iteration capabilities.
Yao Feng: Compared to architecture itself, what's more important is identifying the right technical direction and building the ability to continuously acquire data and feedback.
Zhao Hao: Transformer remains the dominant architecture for now. Future directions worth watching focus more on improving model efficiency, and continued exploration of video generation, multi-scale modeling, and motion capabilities.

Beyond models, technology will also move from labs into industry: Which industries will robots create value in first?
Multiple technical practitioners believe that logistics and pick-and-place operations are foundational action primitives for downstream complex movements, while logistics, factory and supermarket sorting, hotel cleaning, hazardous operations, and other scenarios are seen as having significant application potential.
Additionally, on the question of model-data closed loops, industrial scenario data and models remain difficult to reuse, pre-trained models have limited out-of-the-box capability, and real-world scenario data struggles to effectively feed back into foundation models. Participants agreed that embodied intelligence remains in the early stages of long-term development, and industry needs to continuously advance technology deployment, validating and iterating model capabilities in real scenarios.
This is a discussion without standard answers, and one that continues. Some participants suggested drawing on the development path of autonomous driving, starting from scenarios like logistics where data closed loops can be formed.

Researchers are continuously exploring embodied intelligence technology at every stage of innovation, development, and deployment, and future directions remain wide open.
BlueRun Ventures has been investing in robotics for a decade, and in recent years has continued following the embodied intelligence track, making systematic, deep early-stage bets from multiple technical layers, including AgiBot, Galaxy Universal, Zhijian Power, Tashi Zhihang, Hillbot, Lingchu Intelligence, OriginFlow, and others.
This is why BlueRun hosted this "pure tech" gathering. We believe that real projects and thinking from frontline researchers and technical practitioners matter greatly at this moment. BlueRun will continue following discussions in embodied intelligence technology, hoping to understand the latest directions in technical development, see the possibilities of different technical routes, and together with more researchers and entrepreneurs, observe technical evolution through open discussion, seek long-term accumulation and deep understanding of underlying technology, and explore the key variables in embodied intelligence development. And we will accompany outstanding teams in pushing technological innovation into the real world, finding value entry points and real productivity.


BlueRun's Jui Chan in conversation with BAAI's Wang Zhongyuan, Galaxy Universal's Wang He, and ModelBest's Li Dahai: Long-term value and the next curve in the large model era
BlueRun in Dialogue with AgiBot, Galaxy Universal, Tashi Zhihang, Lingchu Intelligence, Hillbot — Embodied Intelligence: The "Breaking" and "Making" at Dawn
BlueRun AI Annual Outlook: The Survival, Evolution, and Narrative of China's AI Entrepreneurs | Booming Talk


