Astribot's Self-Developed Lumo-1 Model: Enabling Robots to Unite Mind and Hand, Entering the Era of Reasoning-Action Loops | DaoTong Family
On December 11, 2025, **Datong Family company Astribot released LUMO-1, a next-generation embodied reasoning VLA model that successfully established a new paradigm for the robot "reasoning-action" closed loop**. Leveraging core technologies including a three-stage training architecture and structured reasoning, the model comprehensively outperformed mainstream baselines on critical tasks such as long-horizon multi-step manipulation and out-of-distribution generalization, enabling responses without programming.

Background
On December 11, 2025, Astribot, a Datong Family portfolio company, released LUMO-1, a next-generation embodied reasoning VLA model that successfully establishes a new paradigm for the robot "reasoning-action" closed loop. Leveraging core technologies including a three-stage training architecture and structured reasoning, the model surpasses mainstream baselines across critical tasks such as long-horizon multi-step manipulation and out-of-distribution generalization — responding to complex instructions without any programming.
From understanding abstract semantics to dexterously executing fine-grained operations, Lumo-1 upgrades robots from "mechanical mimicry" to "autonomous decision-making," marking a key step toward deploying embodied intelligence in real-world scenarios.
"
Making a robot heat bread

Even though it's never seen this particular loaf, the robot reasons to identify it, infers that heating means using a microwave, and works out the sequence: open door, pick up, place inside, close door, turn knob, wait, remove... No programming needed — the entire long-horizon sequence is completed through pure reasoning!
"
Organizing stationery

Quickly locating stationery amid a cluttered desktop, with fine-grained handling of items in different shapes, materials, and sizes ⚡️
"
Putting a Coke into the blue plate

It even reasons to use its left arm first, but when encountering an obstacle, switches to the right hand for faster retrieval 👏👏👏
Making robots reason like humans so they can act more like humans.
From walking, dancing, to backflips, motion imitation has taught robots how to move. But when it comes to complex operations like carrying plates, sorting fruit, or heating food, robots can't just imitate — they need deep decision-making — recognizing complex environments, understanding the task intent behind why to do something, then translating that into coherent physical actions of how to do it.
From human-like operation to human-like intelligence, embodied intelligence is gradually entering the "mind-hand unity" era of reasoning-action.
Astribot Lumo-1 was born for this moment!
This is an in-house developed end-to-end vision-language-action whole-body VLA model. Through embodied VLM, cross-embodiment joint training, reasoning-action real-robot training, and reinforcement learning calibration alignment — combined with high-quality real-robot training on the cable-driven robot S1 — it translates the "deep mind" of large models into fluid whole-body operation.
Lumo means light and inspiration in Latin. The hope is that it becomes a beam of light for embodied intelligence: enabling robots to "understand what you say," "know why," and then "decide for themselves how to do it."
Project page: www.astribot.com/research/Lumo1
Technical report: https://arxiv.org/pdf/2512.08580
Lumo-1 demonstrates powerful manipulation intelligence and generalization capabilities, surpassing advanced models including π0 and π0.5 across all three core task categories: multi-step long-horizon tasks, fine-grained dexterous manipulation, and generalizable pick-and-place. Its advantages are especially pronounced in out-of-distribution (OOD) scenarios with unseen objects, environments, and instructions, as well as in abstract, ambiguous commands requiring extended reasoning.

Figure: General pick-and-place benchmark results

Figure: Long-horizon and dexterous manipulation task comparison results
Three Key Features Driving "Mind-Hand Unity"

In Lumo-1, 1) Spatial Action Tokenizer (SAT) transforms action trajectories into reusable, composable "action vocabulary," enabling the robot to compose actions like writing sentences, or to reuse, interpret, and predict actions. Technically, SAT compresses continuous action trajectories into minimal path points, and clusters rotational/translational incremental actions into compact tokens. This preserves action-space meaning while reducing irrelevant noise introduced during data collection, proving more compact and stable than FAST and bucketing methods.

Through 2) Structured Reasoning, the robot's brain no longer memorizes trajectories but instead forms structured reasoning chains that explain actions — moving from "executing motion" to "executing ideas." The model performs abstract reasoning around goals, sub-task decomposition, visual element recognition, and spatial action inference, making the why precede the how. Ultimately, it maps visual understanding to path point prediction, allowing 2D predictions to naturally translate into 3D control, achieving more purposeful, context-aware action generation.

Put items that can be used to draw the ocean into the green plate
Strong reasoning doesn't guarantee successful execution. Lumo-1 adds 3) RL Alignment at the final stage, calibrating and aligning high-level reasoning with low-level action. It designs multi-dimensional reward signals covering vision-action-reasoning consistency, action execution, and reasoning format, then uses a GRPO-based learning scheme to encourage the model to select more accurate, coherent, and physically plausible actions. Experiments show this approach significantly surpasses raw imitation of expert demonstrations in task success rate, action rationality, and generalization.
Three-Stage Training: From VLM Cognition to VLA Intelligence
Lumo-1's training isn't about scaling for scale's sake — it's a carefully designed "intelligence transfer" process.
Stage 1: Embodied VLM. Continued pre-training on curated vision-language data gives the model "embodied semantics" including spatial understanding, planning, and trajectory inference. It surpasses specialized models like RoboBrain-7B and Robix-7B on most of 7 classic embodied reasoning benchmarks.

Figure: The curated dataset aims to strengthen core embodied reasoning capabilities without damaging the pre-trained VLM's general multimodal understanding and reasoning abilities.
Stage 2: Cross-Embodiment Joint Training. Joint training across diverse robots, multi-view trajectories, and VLM data strengthens instruction following, object localization, and spatial reasoning — helping the model begin to understand what actions are and how they relate to instructions and observations.
Stage 3: Real-Robot Reasoning-Action Training (S1 trajectories). Using highly human-like demonstration trajectories from the cable-driven robot Astribot S1, the model undergoes action training with reasoning processes, learning executable real-world action patterns: how to handle objects with bimanual coordination, how to execute long-horizon sequences, and how to translate step-by-step reasoning into trajectories.

Figure: Sample tasks collected on the Astribot S1 robot. These tasks cover a broad range of daily activities, collected with diverse objects, lighting conditions, and environmental scenes. Each task involves complex, long-horizon behaviors that naturally decompose into multiple subtasks, containing diverse primitive action units such as sweeping, peeling, pouring, scrubbing, folding, pressing, and rotating.
Finally, RL calibration alignment closes the entire reasoning-action loop, raising average rewards, reducing error rates, and strengthening real-world generalization.
Lumo-1 Training Results Validate Scaling Law
Diverse data is the critical variable — without augmentation, execution rapidly fails in reality. Diverse prompts, image augmentation, and cross-scene training substantially improve robustness. Beyond data volume, embodied intelligence demands data that is "more like the world."
Lumo-1's training results validate the Scaling Law: under data constraints, the model's loss curve closely follows scaling-law predictions, confirming that Lumo-1 is "scalable" — ready to be further amplified.



ID: daltonventure
Long-press to follow
Recommended Reading


