SWE-bench Verified
Software engineering benchmark: resolve real GitHub issues. Verified split filters for human-validated solvable instances.
contamination: mediummetric: resolve_rate
Comparable set
swe-bench-verified/verified/verified/agent/resolve-rate/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 80.6% | ±0.0180 | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 79.3% | ±0.0184 | third-party | 2026-08-10 |
| GLM-5.2 | 78.7% | ±0.0187 | third-party | 2026-08-10 |
| DeepSeek-V4-Pro | 77.6% | ±0.0190 | third-party | 2026-08-10 |
| Qwen3.7-Max | 77.3% | ±0.0191 | third-party | 2026-08-10 |
| GPT-5.4 | 76.9% | ±0.0192 | third-party | 2026-08-10 |
| Kimi-K2.6 | 76.6% | ±0.0192 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 75.6% | ±0.0195 | third-party | 2026-08-10 |
| GPT-5.3 Codex | 74.8% | ±0.0198 | third-party | 2026-08-10 |
| GLM-5.1 | 74.2% | ±0.0199 | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks