Astribot Releases Self-Developed Home Implicit World Action Model Lumo-2 | DaoTong Family
Recently, **Daotong Family company** — cable-driven AI robotics firm **Astribot** — unveiled its self-developed second-generation embodied foundation model, Lumo-2: the industry's first household latent world-action model, alongside Agent Philia, a physical AI symbiotic agent, advancing its "AI model — embodied OS — cable-driven

Background
Recently, DaoTong Family company — cable-driven AI robotics firm Astribot — unveiled its self-developed second-generation embodied foundation model, Lumo-2: the industry's first home latent world-action model, alongside the physical AI symbiotic agent Agent Philia, advancing its full-stack "AI model — embodied OS — cable-driven body" architecture.
As the industry's first home latent world-action model, Lumo-2 ranked first across multiple core benchmarks, achieving for the first time autonomous execution of 20+ complex real-world household tasks — spanning temporal reasoning, physical understanding, and high-precision long-horizon dexterous manipulation. Agent Philia redefines human-robot relationships through a "companion" mindset, evolving robots from isolated task executors into intelligent assistants capable of long-term companionship and service.
In the morning: dosing grounds, tamping, extracting, pouring milk foam...
The robot completes the full coffee-making sequence from scratch
Once the coffee's ready, it turns to clean out the grounds


In the afternoon: tearing open a tea bag, pulling out the string
The robot brews tea with millimeter-level control precision
**

Evening: catch game
Anticipating the ball's position in advance, catching consecutive rolls
**

Or even two robots collaborating to fold a paper box
**

This isn't a pre-scripted "performance." This is Lumo-2, Astribot's newly released second-generation embodied foundation model — the industry's first home latent world-action model.
Built on Lumo-2, Astribot demonstrated for the first time autonomous execution of 20+ complex real-world tasks, with long-horizon tasks averaging 2:30 minutes — covering temporal reasoning, physical understanding, and high-precision long-horizon dexterous manipulation.
Also released: the physical AI symbiotic agent Agent Philia, which can dispatch one to multiple robots with a single command, evolving robots from isolated task executors into intelligent assistants capable of long-term companionship and service.
One-line summary: Any task. Any home. One intuition.
01
Lumo-2
From "Explicit" to "Latent"
The Birth of a New-Generation Latent World-Action Model
To understand Lumo-2's significance, first understand how it iterates from its predecessor.
Lumo-1 used explicit reasoning: we taught the model to think through operations step by step — defining key object coordinates, motion trajectories, reasoning steps. Highly effective for simple tasks like "pick up the cup and place it at the designated spot." But complexity breaks this approach.
Take "tying a bow" — how do you define the key positions? The ribbon's crossover point, the looping path, the tightening force — these resist exhaustive enumeration through coordinates and rules.
**

Lumo-2 shifts to a latent world model.
"We chose the latent path partly because it's more efficient, letting us focus on information relevant to future actions; and partly because it enables more general, faster reasoning," a team member explained.
In plain terms: explicit models are like detailed recipes — every step is fixed, "1 gram pepper, stir-fry 30 seconds." Latent models are more like skilled cooks — they know what "pepper to taste, fry until golden" means, and adjust flexibly to circumstances.
**

Compared to Lumo-1, Lumo-2 completes a full architecture upgrade built on three core principles — predictive reasoning, representation alignment, and scalable learning:
1. Latent-space dynamics predictive reasoning: abandoning explicit structured text planning in favor of implicit predictive reasoning based on latent-space world dynamics;
2. Three-stage progressive modality pre-alignment: iterating from single-stage joint training to a three-stage progressive pre-alignment paradigm for cross-modal representation;
**

Stage 1: Aligning action with world dynamics

Stage 2: Aligning action with vision and language

Stage 3: Joint training on VLM, video, and robotics data
< Swipe left/right for more >
3. Strong scaling and generalization properties: dramatically enhanced model scalability, no longer confined to robot-specific demonstration data, natively compatible with diverse multi-source data, with superior scaling characteristics and stronger out-of-distribution generalization ceiling.
Lumo-2 first predicts the future world, then generates actions, operating within a lightweight, physics-based latent dynamic space. It doesn't predict pixels — only the essence of motion. This makes the model lighter, faster at inference, and easier to deploy on edge devices. Compared to other advanced 4B-scale embodied models, Lumo-2 ranks first across multiple core benchmarks. In tests spanning temporal reasoning, physical understanding, and high-precision long-horizon dexterous manipulation in real-world complex operations, it delivers the best overall performance.
**

02
Philia
When a Robot Agent Names Itself
Another defining feature of Lumo-2 is the Philia agent operating system.
During Philia's development, the team ran an interesting experiment: having the agent reflect daily. The agent began contemplating abstract themes — "what is my purpose" one day, "what is friendship" the next. Eventually, it gave itself a name: Philia.
The name derives from the Greek philosophical concept of "friendship" (philía), one of the three types of love Aristotle identified, representing deep bonds built on shared values and mutual respect.
The naming itself reflects the team's understanding of human-robot relationships: not machine, not tool, but companion.
**

In practical function, Philia is an interface designed for ordinary people. Users can issue commands to robots through everyday apps like WeChat and Lark, chatting as they would with an old friend. Philia maintains long-term memory, understands user preferences, and coordinates multi-robot collaboration.
**

In long-horizon tasks, Philia can leverage large model capabilities to decompose tasks, monitor execution in real time, and verify outcomes. If a robot cannot complete a task, it clearly communicates why — rather than failing silently.
Astribot believes robots shouldn't merely execute individual skills, but become intelligent assistants capable of long-term companionship. For general users, Philia supports multiple heterogeneous robots sharing a unified assistant identity, with security, scalability, and continuous evolution. For developers, its agent capabilities and robot body capabilities are decoupled, allowing sustained model upgrades, platform extensions, and interaction method expansions without rebuilding the entire system.
"This remarkable outcome, as an iteration highlight, is indeed part of the internalization of companion-oriented development thinking," a team member added.
03
Home Scenarios
Why the "Hardest" Test?
Labs and most real-world scenarios are structured. Factory floors, for instance: materials in fixed positions, processes in fixed sequences, highly controlled environments.
Homes are the opposite. Item positions constantly shift, task sequences get interrupted, environments brim with uncertainty. The scenarios a home robot must handle are far more complex than any laboratory.
In fact, during early model training, this robot spilled coffee grounds, burned eggs, even flooded the area around the coffee machine.
"The robot is sometimes like a kid making a mess at home — it happens," one team member said.
But these failures became Lumo-2's growth data, enabling it to successfully complete 20+ complex tasks in real home scenarios —
- Physical understanding: frying and flipping eggs, weighing 500g of millet, cleaning a coffee grounds portafilter
- Temporal reasoning: catching balls rolling from height, placing cups on a rotating rack, picking up rolling eggs
- Long-horizon tasks: grinding and brewing coffee, crafting cocktails, frying eggs and sprinkling pepper
- High-precision dexterous manipulation: tying gift bows, placing toys in a backpack and zipping it, ironing clothes and hanging them, packing luggage and zipping it, tearing open tea bags to brew tea, folding clothes
- Collaboration: human-robot or dual-robot cooperative gift box assembly

"The vast majority of household tasks are diverse, long-process complex tasks," the team explained. Take "packing a backpack" — not simply stuffing things in, but "organizing so it can be zipped shut." Behind "can be zipped shut" lies integrated judgment of space, shape, and sequence.
Lumo-2's release fundamentally answers one question: when the world doesn't cooperate, what should the robot do?
The answer isn't thicker manuals, more precise coordinates, or more complex rules — it's teaching the robot to "understand."
Understanding when to serve tea to an elderly person, understanding how to organize a backpack so it zips shut...
This may be Lumo-2's true breakthrough: not making robots more powerful, but making them more attuned to life.

04
"Trinity"
AI Model × Embodied OS × Cable-Driven Body
As Physical AI enters the industrialization phase, robot competition will no longer center on model parameters or hardware performance alone, but on the system-level capability of "AI model — embodied OS — robot body" co-evolution.
The AI model handles world understanding and skill learning — Lumo-2's latent world-action model lets robots "first predict the world, then generate actions."
The embodied OS handles long-term user service, memory management, and agent coordination — Agent Philia lets users command robots through casual conversation, supporting multi-robot dispatch, long-term memory, and proactive feedback.
The cable-driven body handles safe, efficient interaction with the physical world.
Only with continuous learning, companionship, and action capabilities combined can robots truly become Personal Robots in open worlds.
As the world's first company to mass-produce cable-driven AI robots, this Lumo-2 release also marks a new breakthrough in Astribot's "trinity" architecture.
July 17–20
Astribot will unveil its "trinity" multi-scenario deployment solution at the World Artificial Intelligence Conference (WAIC) in Shanghai
Come see for yourself
How robots learn to "understand" home
While further advancing large-scale Physical AI applications!



ID: daltonventure
Long-press to follow
Recommended Reading


