Product

SWE-bench Pro

SWE-Bench Pro

SWE-bench Pro is a software engineering benchmark used to evaluate AI models' coding capabilities. As described in a late-2025 ZhenFund interview, it sits at the frontier of how the industry measures model progress: Gemini 3 Pro scored roughly 78 on it, while Opus 4.5 reached the low 80s — a narrow numerical gap that, as Yusen Dai noted, can obscure large differences in real-world user experience . In early 2026, the Chinese AI company VAST reported its M2.7 model achieved 56.22% on SWE-Pro, calling that "close to current first-tier model levels," suggesting the benchmark has become a standard reference point for frontier coding performance . The benchmark itself appears to be an evolved or professional-tier version of the original SWE-bench .

AI-generated — may contain errors, please verify.

SWE-bench ProProduct
SWE-Bench Pro
No graph yet
Mentioned in 4 articles

Coverage