SimpleQA Verified
Short-answer factual recall on questions with a single unambiguous answer, after a cleaning pass that removed disputed or outdated items.
contamination: mediummetric: accuracy
Comparable set
simpleqa-verified/verified/verified/0shot/exact-match/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview) | 77.3% | ±0.0133 | third-party | 2026-08-10 |
| GPT-5.6 Sol | 71.6% | ±0.0143 | third-party | 2026-08-10 |
| Gemini 3.6 Flash | 68.7% | ±0.0147 | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 68.4% | ±0.0147 | third-party | 2026-08-10 |
| Claude Fable 5 | 68.3% | ±0.0147 | third-party | 2026-08-10 |
| GPT-5.5 | 64.5% | ±0.0151 | third-party | 2026-08-10 |
| Qwen3.7-Max | 58.5% | ±0.0156 | third-party | 2026-08-10 |
| DeepSeek-V4-Pro | 57.0% | ±0.0157 | third-party | 2026-08-10 |
| Claude Opus 5 | 56.7% | ±0.0157 | third-party | 2026-08-10 |
| Grok 4.5 | 53.5% | ±0.0158 | third-party | 2026-08-10 |
| Qwen3-235B-A22B-Instruct-2507 | 50.1% | ±0.0158 | third-party | 2026-08-10 |
| GPT-5.4 | 47.8% | ±0.0160 | third-party | 2026-08-10 |
| Qwen3.8-Max | 46.3% | ±0.0158 | third-party | 2026-08-10 |
| GPT-5.6 Terra | 43.1% | ±0.0157 | third-party | 2026-08-10 |
| Kimi-K3 | 42.7% | unknown | third-party | 2026-08-10 |
| GPT-5.6 Luna | 41.7% | ±0.0156 | third-party | 2026-08-10 |
| Claude Opus 4.8 | 39.5% | ±0.0155 | third-party | 2026-08-10 |
| Kimi-K2.7-Code | 39.2% | ±0.0154 | third-party | 2026-08-10 |
| Kimi-K2.6 | 38.7% | ±0.0154 | third-party | 2026-08-10 |
| GLM-5.2 | 38.1% | ±0.0154 | third-party | 2026-08-10 |
| Grok 4.3 | 38.0% | ±0.0154 | third-party | 2026-08-10 |
| GLM-5.1 | 37.3% | ±0.0153 | third-party | 2026-08-10 |
| DeepSeek-V4-Flash-0731 | 34.7% | unknown | third-party | 2026-08-10 |
| GLM-4.7 | 31.5% | unknown | third-party | 2026-08-10 |
| GPT-5.4 mini | 28.6% | ±0.0143 | third-party | 2026-08-10 |
| DeepSeek-R1-0528 | 27.4% | unknown | third-party | 2026-08-10 |
| Claude Sonnet 5 | 25.0% | ±0.0137 | third-party | 2026-08-10 |
| gpt-oss-120b | 13.9% | unknown | third-party | 2026-08-10 |
| GPT-5.4 nano | 12.0% | ±0.0103 | third-party | 2026-08-10 |
| gemma-4-31B-it | 9.6% | ±0.0094 | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks