SWE-bench Verified

Software engineering benchmark: resolve real GitHub issues. Verified split filters for human-validated solvable instances.

contamination: mediummetric: resolve_rate

Comparable set

swe-bench-verified/verified/verified/agent/resolve-rate/inspect

ModelScoreUncertaintyRun byRetrieved
GPT-5.580.6%±0.0180third-party2026-08-10
Gemini 3.5 Flash79.3%±0.0184third-party2026-08-10
GLM-5.278.7%±0.0187third-party2026-08-10
DeepSeek-V4-Pro77.6%±0.0190third-party2026-08-10
Qwen3.7-Max77.3%±0.0191third-party2026-08-10
GPT-5.476.9%±0.0192third-party2026-08-10
Kimi-K2.676.6%±0.0192third-party2026-08-10
Gemini 3.1 Pro (preview)75.6%±0.0195third-party2026-08-10
GPT-5.3 Codex74.8%±0.0198third-party2026-08-10
GLM-5.174.2%±0.0199third-party2026-08-10

Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks