SWE-bench Pro
SWE-Bench Pro
SWE-bench Pro is a software engineering benchmark used to evaluate AI models' coding capabilities. As described in a late-2025 ZhenFund interview, it sits at the frontier of how the industry measures model progress: Gemini 3 Pro scored roughly 78 on it, while Opus 4.5 reached the low 80s — a narrow numerical gap that, as Yusen Dai noted, can obscure large differences in real-world user experience . In early 2026, the Chinese AI company VAST reported its M2.7 model achieved 56.22% on SWE-Pro, calling that "close to current first-tier model levels," suggesting the benchmark has become a standard reference point for frontier coding performance . The benchmark itself appears to be an evolved or professional-tier version of the original SWE-bench .
AI-generated — may contain errors, please verify.
Coverage
MiniMax Quietly Launched M3 and MiniMax Code — Here's What We Found After Testing Them
Cutting-edge coding capabilities, a 1M context window, and native multimodality.
Moonshot AI K2.6 Released and Open-Sourced: Major Advances in Coding, Long-Horizon Tasks, and Agent Swarms | BlueRun Ventures Portfolio Headlines
Minor Version, Major Upgrade
Moonshot AI's K2.6 Model Is Here: Major Leaps in Long-Range Coding and Agent Swarm Capabilities!
Talk is cheap. Show me the code.
Moonshot AI K2.6: The Leap in Long-Horizon Execution and Agent Swarm Capabilities
Thirteen hours of continuous execution



