Product

Vision-Language-Action Model

VLA模型

A Vision-Language-Action (VLA) model is a robotics architecture that extends visual-language models (VLMs) to directly output robot control signals rather than text descriptions . It takes multimodal inputs—camera images, text instructions, and robot state data like joint angles—and generates precise, high-frequency action commands for physical execution .

The field emerged after VLMs matured, with Google's RT-2 (2023) among the first notable implementations, adapting PaLI-X and PaLM-E models by encoding robot actions as text tokens for end-to-end training . By late 2024, industry efforts like Gemini for Robotics and NVIDIA GR00T followed . Recent technical evolution has moved toward decoupling semantic understanding from control dynamics: the CLAP framework, for instance, trains separate VLA models for task planning and precise action execution, using a "latent action" space to bridge human video demonstrations with robot trajectories . AgiBot's GO-1 (2025) pushed further with a ViLLA architecture that adds an implicit latent planner between VLM perception and a diffusion-based action expert .

VLA remains contested. Proponents see it as the practical path to "see and do" robotics for structured tasks like warehouse manipulation . Critics, notably Yann LeCun, argue that VLA's reliance on behavior cloning from massive demonstration datasets limits generalization, and that its end-to-end black-box design lacks explicit planning and consequence prediction—capabilities he associates with world-model approaches . A 2026 assessment from 十字路口Crossing characterized VLA as "top-heavy" on language parameters, strong at conceptual transfer but weak at genuine physical understanding .

AI-generated — may contain errors, please verify.

Vision-Language-Action ModelProduct
VLA模型
渲染中…
Mentioned in 7 articles

Coverage