Vision-Language-Action Model
VLA模型
A Vision-Language-Action (VLA) model is a robotics architecture that extends visual-language models (VLMs) to directly output robot control signals rather than text descriptions . It takes multimodal inputs—camera images, text instructions, and robot state data like joint angles—and generates precise, high-frequency action commands for physical execution .
The field emerged after VLMs matured, with Google's RT-2 (2023) among the first notable implementations, adapting PaLI-X and PaLM-E models by encoding robot actions as text tokens for end-to-end training . By late 2024, industry efforts like Gemini for Robotics and NVIDIA GR00T followed . Recent technical evolution has moved toward decoupling semantic understanding from control dynamics: the CLAP framework, for instance, trains separate VLA models for task planning and precise action execution, using a "latent action" space to bridge human video demonstrations with robot trajectories . AgiBot's GO-1 (2025) pushed further with a ViLLA architecture that adds an implicit latent planner between VLM perception and a diffusion-based action expert .
VLA remains contested. Proponents see it as the practical path to "see and do" robotics for structured tasks like warehouse manipulation . Critics, notably Yann LeCun, argue that VLA's reliance on behavior cloning from massive demonstration datasets limits generalization, and that its end-to-end black-box design lacks explicit planning and consequence prediction—capabilities he associates with world-model approaches . A 2026 assessment from 十字路口Crossing characterized VLA as "top-heavy" on language parameters, strong at conceptual transfer but weak at genuine physical understanding .
AI-generated — may contain errors, please verify.
Coverage
Every company in the embodied intelligence space could be solid — the key is whether they've figured out who they want to become | Linear Voice
The embodied data foundation needs to be at the scale of 10 million hours or more.
DeepRoute's Self-Developed VLA Goes Live, First Batch of Co-Developed Mass-Production Vehicles Coming Soon | Yunqi Capital
Entering the Era of "Defensive Driving"
MaHui Entrepreneur | Xingxing Wang: The Biggest Barrier to Mass Humanoid Robot Deployment Is Embodied Artificial Intelligence
MaHui portfolio company Unitree founder and CEO Xingxing Wang shared his latest observations and insights on the global robotics industry at World Robot Conference 2025.
Yunqi Capital | DeepRoute.ai x Volcano Engine, Partnering to Accelerate Agent Deployment in Vehicles
From Driving Machine to Intelligent Agent: The Evolution of the Car
Yunqi Capital | Angel-round project *DeepRoute* raises $100 million in Series C1, accelerating mass production of advanced intelligent driving
will also explore large-scale Robotaxi operations
Yunqi Capital Quarterly | We Are Stardust, Imagining the Universe
New Practices, New Breakthroughs, New Thinking for Summer and Fall
ChatGPT for Robots: Large Models Enter the Physical World, DeepMind's Breakthrough | BlueRun Ventures Share
Is Embodied Intelligence Close?






