Surpassing Pi 0.5! Spirit AI Open-Sources Spirit v1.5, Topping Key Benchmarks | Oasis Vitality
Counselor Vitality

On January 12, Spirit AI, a seed-stage portfolio company of Oasis Capital, announced that its self-developed embodied intelligence model Spirit v1.5 ranked first overall in the RoboChallenge benchmark, surpassing Pi 0.5 in both task scores and success rate.

To verify that its ranking came from a genuinely self-developed model, Spirit AI simultaneously open-sourced relevant information about Spirit v1.5, inviting independent scrutiny from the public and research community. Through this approach, researchers can not only reproduce the benchmark results but also use Spirit v1.5 as a foundational embodied intelligence model for further research and innovation.
RoboChallenge is a standardized evaluation framework newly established in 2025, jointly launched by Dexmal, Hugging Face, and other institutions, focused on cross-platform validation of embodied intelligence models. As a key benchmark in the current embodied intelligence field that emphasizes real robot execution capabilities, its evaluation tasks cover complex instruction understanding, multi-step operation planning, and cross-scenario execution stability, among other dimensions.
Spirit v1.5's first-place finish on this platform demonstrates its comprehensive capabilities in general robot tasks and real-world execution scenarios.

RoboChallenge Evaluation Performance Overview
Looking at the evaluation results, Spirit v1.5 maintained high success rates across multiple tasks, showing particularly stable performance in dimensions such as multi-task continuous execution, complex instruction decomposition, and cross-morphology transfer. As of the latest evaluation cycle, its overall score exceeded previously leading models such as Pi 0.5, placing it at the top of the leaderboard.
RoboChallenge's scoring system evaluates not only whether tasks are completed but also the model's execution process, including spatial localization, occlusion handling, long-horizon stability, and transfer efficiency when facing novel tasks. This evaluation approach imposes higher demands on a model's generalization, stability, and execution accuracy — and more closely approximates real-world robotics applications.

Technical Architecture and Key Methods
In terms of model architecture, Spirit v1.5 adopts a Vision-Language-Action (VLA) unified modeling framework, integrating visual perception, language understanding, and action generation within a single decision-making pipeline. This reduces information loss from chaining multiple modules and improves overall stability in long-horizon tasks.
On the training methodology front, a core characteristic of Spirit v1.5 is that it does not rely on highly curated "clean" demonstration data. In a technical blog post, Spirit AI argued that overly scripted, controlled-environment data collection, while beneficial for rapid model convergence, ultimately limits generalization in the real world.

Accordingly, Spirit v1.5 introduced an open-ended, diverse data collection paradigm during pre-training. Data collection was no longer strictly bound to task scripts; instead, it was oriented around "accomplishing meaningful goals," allowing natural chaining of multiple sub-tasks and atomic skills during operation. This approach exposed the model during training to complexity closer to the real world, including occlusions, failure recovery, and natural transitions between tasks.
Related ablation experiments showed that, at equivalent data scales, models pre-trained on diverse data achieved significantly higher transfer efficiency on novel tasks compared to those trained on traditional demonstration data, requiring markedly fewer computational resources to reach the same performance level. This finding helps explain Spirit v1.5's stable performance on RoboChallenge's multi-morphology, unseen-task evaluations.

Open-Source Strategy and Community Significance
Alongside its benchmark results, Spirit AI chose to simultaneously open-source Spirit v1.5's base model weights, inference code, and usage examples. Through this approach, the community can not only verify the model's performance but also extend it as a foundational model for embodied intelligence research.
At a time when embodied intelligence research still heavily depends on a handful of technical approaches, Spirit v1.5 offers academia and industry an alternative data paradigm and training methodology, helping advance exploration toward more generalizable universal robot models.
(Click "Read Original" to jump to Spirit AI's technical blog post)





