Product

Vision-Language-Action

VLA

Vision-Language-Action (VLA) is a class of multimodal AI models that extends visual-language models (VLMs) into robotics, translating visual perception and language instructions directly into physical action commands rather than text outputs . As described by Oasis Capital, VLA models integrate robot state data—such as joint angles and arm position—alongside camera feeds and text prompts, with the goal of outputting control signals at the high frequency and low latency required for real-time manipulation .

The architecture has evolved from early approaches where LLMs generated control signals directly, to newer designs like π0.7 that route through a dedicated "action expert" module . Google Research's RT-2 was among the first prominent examples, training on web data, robot demonstrations, and other multimodal sources . By late 2024, industry efforts like Gemini for Robotics and NVIDIA GR00T had joined the field .

VLA sits at the center of a technical debate: Yann LeCun has argued that VLA models amount to behavior cloning that fails to generalize beyond training distributions, requiring "massive demonstration samples" and collapsing when faced with novel situations . In contrast, firms like Stardust Intelligence are betting on end-to-end VLA architectures such as Lumo-1, which pairs embodied VLMs with cross-embodiment joint training and reinforcement learning to achieve whole-body manipulation . FreeS Fund's Li Feng, meanwhile, cautions that the visual-to-language-to-action chain involves "complex perception, cognition, and execution processes" that data accumulation alone cannot bridge .

VLA is increasingly framed not as competing with world models but as complementary—"see and do" for structured, short-horizon tasks versus "think first, then do" for uncertain, long-horizon planning—with recent research like World VLA exploring hybrid architectures .

AI-generated — may contain errors, please verify.

Vision-Language-ActionProduct
VLA
渲染中…
Mentioned in 21 articles

Coverage