MiniMax M2.5 Released, Accelerating the Era of Agents for Everyone | Oasis Capital Vitality
$1/hour, the real-world work champion

In AI, the speed of evolution is often one of the hardest metrics of competitiveness.
Oasis Capital portfolio company MiniMax today officially released M2.5, the latest iteration of its M2 series, signaling that large models are evolving deeply toward "native Agent" capabilities. It has not only set new SOTA records in high-end scenarios like programming, search, and office productivity, but achieved breakthroughs across two dimensions:
- Thinking like an architect: M2.5 is no longer just a "typewriter" that writes code — it now possesses end-to-end design capabilities spanning functional decomposition to system testing. At MiniMax, code generated by the model already accounts for 80% of all code.
- Making Agents economically viable: The operating cost of complex Agents has long been an industry pain point. M2.5 delivers a 37% improvement in end-to-end execution speed, while cutting continuous operation costs to one-tenth of mainstream models. This means "$10,000 supporting four Agents working around the clock, year-round" is becoming reality.
From empowering R&D and sales to deepening into specialized professional domains like finance and law, M2.5 is accelerating the arrival of the "Agent for everyone" era. Below is our breakdown of MiniMax M2.5's core upgrades.
Enjoy
The M2 series is iterating at the fastest pace in the industry. Today, we introduce MiniMax M2.5
- M2.5 has reached or set new industry SOTA benchmarks across productivity scenarios including programming, tool use, search, and office work — for example, SWE-Bench Verified (80.2%), Multi-SWE-Bench (51.3%), and BrowseComp (76.3%);
- M2.5 optimizes the model's ability to decompose complex tasks and reduces token consumption during reasoning, enabling faster completion of complex Agentic tasks. On SWE-Bench Verified tests, M2.5 completes tasks 37% faster than the previous version M2.1;
- M2.5 makes running complex Agents indefinitely economically viable. At 100 tokens per second output, M2.5 costs just $1 per hour of continuous work; at 50 tokens per second, only $0.30.

MiniMax has already been among the first to benefit from M2.5's capabilities. In real internal business scenarios, 30% of overall tasks are autonomously completed by M2.5, covering functions including R&D, product, sales, HR, and finance — with penetration continuing to rise. Programming has been a standout: M2.5-generated code now accounts for 80% of newly submitted code.
We expect M2.5 to accelerate the arrival of the Agent-for-everyone era.

Programming: Thinking and Building Like an Architect
On core programming benchmarks, M2.5 has improved significantly over the previous generation, reaching levels comparable to the Claude Opus series. On the multilingual task Multi-SWE-Bench, M2.5 achieved first place.

M2.5 has developed the ability to "think and build like an architect." For instance, the model has evolved native Spec behavior: before writing any code, it proactively decomposes functionality, structure, and UI design from an architect's perspective, completing comprehensive upfront planning.
M2.5 was trained on more than 10 languages (including Go, C, C++, TypeScript, Rust, Kotlin, Python, Java, JavaScript, PHP, Lua, Dart, Ruby) across hundreds of thousands of real environments. Beyond bug-fix scenarios, it delivers reliable performance across the full spectrum of complex system development — from 0-to-1 system design and environment setup, 1-to-10 system development, 10-to-90 feature iteration, to 90-to-100 comprehensive code review and system testing. It handles full-stack projects across Web, Android, iOS, Windows, and Mac platforms, encompassing server-side APIs, business logic, databases, and more — not just "frontend webpage demos."
To measure these capabilities, we upgraded our VIBE benchmark to a more complex and challenging Pro version: significantly increasing task complexity, domain coverage, and evaluation accuracy. Overall, M2.5 performs on par with Opus 4.5.

We examined the model's generalization across different scaffolding environments. We tested performance on the SWE-Bench Verified benchmark across various programming scaffolds. On Droid, M2.5 achieved a 79.7% pass rate, surpassing M2.1's 71.3% and Opus 4.6's 78.9%; on OpenCode, M2.5 scored 76.1%, exceeding M2.1's 72.0% and Opus 4.6's 75.9%.

Search and Tool Use: Solving Problems More Efficiently

Search and tool use are prerequisites for models to autonomously handle complex tasks. On benchmarks including BrowseComp and Wide Search, M2.5 has reached top-tier industry levels. The model's generalization capabilities have also improved — M2.5 demonstrates more stable performance when facing unfamiliar scaffolding environments.
In real-world expert search tasks, using a search engine is only a small part; most of the work involves deep exploration within professional websites. To this end, we built RISE (Realistic Interactive Search Evaluation) to measure model performance on realistic professional search tasks. Results show that M2.5 excels at expert-level search tasks in real-world settings.
Compared to its predecessor, M2.5 also demonstrates greater decision-making maturity when handling complex tasks: it has learned to solve problems with more precise search rounds and better token efficiency. For example, across BrowseComp, Wide Search, and RISE tasks, M2.5 achieves better results with fewer rounds, saving approximately 20% in round consumption compared to M2.1. This indicates the model is no longer just "getting the answer right" — it's converging on results through leaner paths.

Office Scenarios: Delivering Professional Output Directly
We considered how to produce truly deliverable work products in office scenarios. To this end, we partnered deeply with senior practitioners in finance, law, social sciences, and other fields — having them propose requirements, provide feedback, participate in standard definition, and directly build datasets, bringing industry tacit knowledge into the model's training pipeline. On this foundation, M2.5 has achieved significant capability improvements in advanced office scenarios including Word, PowerPoint, and Excel financial modeling. At the evaluation level, we built an internal Cowork Agent evaluation framework (GDPval-MM), using pairwise comparison to assess model delivery quality and trajectory professionalism, while monitoring full-process token costs to estimate actual productivity benefits. Against mainstream models, it achieved a 59.0% average win rate.



Fast Reasoning on Complex Tasks
We always want Agents to complete complex tasks as quickly as possible. This depends on the model's ability to decompose complex tasks, its token efficiency, and its inference speed. Our model already provides 100 TPS inference speed — nearly double that of current mainstream models. Meanwhile, through reinforcement learning, we focused on optimizing the model's complex task decomposition ability and token consumption during reasoning. These three factors combined give M2.5 significant advantages in both time and cost for completing complex tasks.
For example, when running the SWE-Bench Verified benchmark, M2.5 consumed an average of 3.52M tokens per task. By comparison, M2.1 would consume 3.72M tokens. Meanwhile, thanks to improvements in parallel tool calling and other capabilities, end-to-end execution dropped from an average of 31.3 minutes to 22.8 minutes — a 37% speed improvement. This timing is essentially on par with Claude Opus 4.6's 22.9 minutes.

Continuous Operation Without Cost Burden
We designed the M2 series with the intention of enabling complex Agents to run without cost constraints. As our capabilities have continued to improve, we believe M2.5 has nearly achieved this goal. M2.5 offers two versions with identical capabilities but different speeds and prices: a fast version at around 100 TPS, costing just $0.30 per million input tokens and $2.40 per million output tokens; and a 50 TPS version with output prices half as low. Referenced by output pricing, the 50 TPS version costs one-tenth to one-twentieth of models like Opus, Gemini 3 Pro, and GPT-5.
At 100 tokens per second output, continuous work for one hour costs only $1; at 50 tokens per second, just $0.30. In other words, $10,000 can support four Agents working continuously for a full year. M2.5 makes building and operating Agents economically almost unlimited. For the M2 series, the only remaining question is the pace of model capability improvement.

The Fastest Improvement Rate in the Industry
Over the past 108 days, we have successively released M2, M2.1, and M2.5 — with model improvement exceeding our original expectations. For instance, on SWE-Bench Verified, the most representative benchmark in programming, the M2 series has maintained the fastest improvement rate in the industry compared to model families like Claude, GPT, and Gemini.


Native Agent RL Framework
We believe the core driver of these advances is large-scale reinforcement learning. It has significantly improved model capabilities and generalization across scaffolding and environments. Through co-design of the Agent RL framework, algorithms, reward design, and engineering optimization, we support efficient optimization for arbitrary Agent scaffolding and environments. We validated near-linear scaling of model capabilities with compute and task volume through large-scale training on hundreds of thousands of Agent scaffolding environments, including extensive real internal company tasks.
Forge — Native Agent RL Framework:
Forge, as a native Agent RL framework, introduces an intermediate layer that fully decouples the underlying training and inference engine from Agents, supporting arbitrary Agent integration and enabling us to optimize model generalization across Agent scaffolding and tools. To improve system throughput, we optimized asynchronous scheduling strategies to balance system throughput and sample off-policyness, and designed a tree-structured merge training sample strategy, achieving approximately 40x training speedup.

Agent RL Algorithms and Reward Design:
At the algorithm level, we continued using our CISPO algorithm proposed earlier this year to ensure stability in large-scale training of MoE models. To address the credit assignment challenge posed by long contexts in Agent scenarios, we introduced process reward mechanisms (Process Reward) for full-chain monitoring of completion quality. Additionally, to deeply align with user experience, we directly estimate task time consumption in real environments and use it as reward, achieving better balance between model performance and response speed.

We will share more about RL scaling and the Agent RL framework Forge in a subsequent technical blog post.

The Best Agentic Experience
M2.5 is now fully live in MiniMax Agent, delivering the best Agentic experience.
We distilled core information-processing capabilities into standardized Office Skills, deeply integrated into the Agent. In Omni (MAX) mode, when handling tasks like Word formatting, PowerPoint editing, and Excel calculations, MiniMax Agent automatically loads the corresponding Office Skills based on file type, improving output quality. Additionally, users can combine Office Skills with domain-specific industry expertise to create reusable Experts tailored to particular task scenarios.
Take industry research as an example: when mature research framework SOPs are fused with Word Skills, the Agent can strictly follow the established framework to automatically pull data, organize analytical logic, and output properly formatted research reports — rather than merely generating rough text. In financial modeling scenarios, when institutional modeling standards are combined with Excel Skills, the Agent can follow specific risk control logic and calculation standards to automatically generate and validate complex financial models, instead of just outputting a simple table.
To date, users have built over 10,000 Experts on MiniMax Agent, with numbers growing rapidly. MiniMax has also built multiple deeply optimized, ready-to-use Expert sets on MiniMax Agent for high-frequency scenarios including office work, finance, and programming.
M2.5 is now fully live across all MiniMax products:
- MiniMax Agent: agent.minimaxi.com
- M2.5 API access: platform.minimaxi.com/docs/guides/text-generation
- Coding Plan subscription: platform.minimaxi.com/subscribe/coding-plan
MiniMax M2.5 model weights will be open-sourced on HuggingFace, supporting local deployment.





