Vision-Language-Action
VLA
Vision-Language-Action (VLA) is a class of multimodal AI models that extends visual-language models (VLMs) into robotics, translating visual perception and language instructions directly into physical action commands rather than text outputs . As described by Oasis Capital, VLA models integrate robot state data—such as joint angles and arm position—alongside camera feeds and text prompts, with the goal of outputting control signals at the high frequency and low latency required for real-time manipulation .
The architecture has evolved from early approaches where LLMs generated control signals directly, to newer designs like π0.7 that route through a dedicated "action expert" module . Google Research's RT-2 was among the first prominent examples, training on web data, robot demonstrations, and other multimodal sources . By late 2024, industry efforts like Gemini for Robotics and NVIDIA GR00T had joined the field .
VLA sits at the center of a technical debate: Yann LeCun has argued that VLA models amount to behavior cloning that fails to generalize beyond training distributions, requiring "massive demonstration samples" and collapsing when faced with novel situations . In contrast, firms like Stardust Intelligence are betting on end-to-end VLA architectures such as Lumo-1, which pairs embodied VLMs with cross-embodiment joint training and reinforcement learning to achieve whole-body manipulation . FreeS Fund's Li Feng, meanwhile, cautions that the visual-to-language-to-action chain involves "complex perception, cognition, and execution processes" that data accumulation alone cannot bridge .
VLA is increasingly framed not as competing with world models but as complementary—"see and do" for structured, short-horizon tasks versus "think first, then do" for uncertain, long-horizon planning—with recent research like World VLA exploring hybrid architectures .
AI-generated — may contain errors, please verify.
Coverage
Which Path Leads to the "World Model" Endgame? | A Conversation with Biwei Huang: Founder of Aether AI
Causal World Models: A New Paradigm for AI
Is the $1 Billion "World Model" Bet by a Turing Award Winner AI's Next Decade? (Part 2)
True intelligence doesn't begin with language, but with the world.
Which Path Leads to the "World Model" Endgame? | A Conversation with Biwei Huang, Founder of Aether AI
It's 2026, and we're still evaluating World Models the way Edison tested filaments.
A Serious Discussion on World Models
A 20-Minute Talk by a NVIDIA Scientist: The Endgame for Robots, a 2040 Prophecy
From VLA to WAM, NVIDIA Is Rewriting Embodied Artificial Intelligence Using the LLM Playbook
Billions of years ago, living organisms built the first "world model."
How Far Are We From a World Model, Now That We've "Uploaded" a Fruit Fly Brain?
The Data Path Toward Embodied AI's GPT-3 Moment | Yang Gao x Beipo Initiative
"Massive amounts of data will produce intelligence."
Gaoyang: A Thinking Reed, A Machine That Acts | Beipo Initiative
The "OpenAI Moment" for Pharma Labs: HeTan AI Raises 50 Million Yuan, Robot Scientists Step Onto the Lab Bench
The story of robot scientists has only just begun.
Obsession, Ambition, and a Pair of AI Glasses: The Raw Fuel of Elite Product Managers | A Conversation with Li Auto SVP Fan Haoyu
Not content to let the world stay as it is.
DeepRoute's Real-World Road Test: When AI Learns to "Fear," the "Black Box" of Assisted Driving Is Opened | Yunqi Capital
When Cars Begin to Understand the World
Yunqi Capital Quarterly | Upward, the Consistent Answer
New Growth, New Gains
Wang Qian, Invariant Robot: How Far Is Embodied Intelligence's Scaling Law? | Yunqi Capital Doers Series
Embodied Intelligence ≠ Stuffing DeepSeek into a Unitree Robot
Yunqi Capital | Yuanrong Qixing Surpasses 30,000 Units in Mass Production Deliveries for September, Setting Another Record
How VLAs Are Pioneering a New Future for Driving
Is AI the Decisive Factor for Embodied Intelligence? After Raising 300 Million in Six Months, How Is VBot's First Product Shaping Up? | A Conversation with Vbot Co-founder Zhe Lun Zhao
I want to make a product that I can bring home to my mom.
Seeing the Rhythm in Progress | Vital Monthly
Counselor Vitality
LimX Dynamics Secures Strategic Lead Investment from JD.com to Accelerate Embodied AI Commercialization | Oasis Capital Vitality
Counselor Vitality --- *Note: "参赞生命力" appears to be a fragment without surrounding context. If this is a title or heading, it could also be rendered as "Nourishing Vitality" or "Fostering Life Force" depending on the specific context — "参赞" (cānzàn) carries connotations of participation, assistance, or counsel, often used in classical/philosophical contexts (e.g., the Confucian idea of humans "assisting" the transforming and nourishing process of heaven and earth). Please provide more context for a more precise translation.*
Spirit AI Closes 528 Million Yuan Funding Round, Accelerating Full-Scale Iteration | Oasis Capital Vitality
Counselor Vitality
AI-Powered, Accelerating Evolution! Five BlueRun Ventures Portfolio Companies Named to ChinaVenture's "Sharp 100" List
Show sharp edge, charge ahead all the way — or more naturally: **Sharp and unstoppable, full speed ahead.**
The Moon and Sixpence of Embodied Intelligence
Counselor Vitality



















