SimpleQA Verified

Short-answer factual recall on questions with a single unambiguous answer, after a cleaning pass that removed disputed or outdated items.

contamination: mediummetric: accuracy

Comparable set

simpleqa-verified/verified/verified/0shot/exact-match/inspect

ModelScoreUncertaintyRun byRetrieved
Gemini 3.1 Pro (preview)77.3%±0.0133third-party2026-08-10
GPT-5.6 Sol71.6%±0.0143third-party2026-08-10
Gemini 3.6 Flash68.7%±0.0147third-party2026-08-10
Gemini 3.5 Flash68.4%±0.0147third-party2026-08-10
Claude Fable 568.3%±0.0147third-party2026-08-10
GPT-5.564.5%±0.0151third-party2026-08-10
Qwen3.7-Max58.5%±0.0156third-party2026-08-10
DeepSeek-V4-Pro57.0%±0.0157third-party2026-08-10
Claude Opus 556.7%±0.0157third-party2026-08-10
Grok 4.553.5%±0.0158third-party2026-08-10
Qwen3-235B-A22B-Instruct-250750.1%±0.0158third-party2026-08-10
GPT-5.447.8%±0.0160third-party2026-08-10
Qwen3.8-Max46.3%±0.0158third-party2026-08-10
GPT-5.6 Terra43.1%±0.0157third-party2026-08-10
Kimi-K342.7%unknownthird-party2026-08-10
GPT-5.6 Luna41.7%±0.0156third-party2026-08-10
Claude Opus 4.839.5%±0.0155third-party2026-08-10
Kimi-K2.7-Code39.2%±0.0154third-party2026-08-10
Kimi-K2.638.7%±0.0154third-party2026-08-10
GLM-5.238.1%±0.0154third-party2026-08-10
Grok 4.338.0%±0.0154third-party2026-08-10
GLM-5.137.3%±0.0153third-party2026-08-10
DeepSeek-V4-Flash-073134.7%unknownthird-party2026-08-10
GLM-4.731.5%unknownthird-party2026-08-10
GPT-5.4 mini28.6%±0.0143third-party2026-08-10
DeepSeek-R1-052827.4%unknownthird-party2026-08-10
Claude Sonnet 525.0%±0.0137third-party2026-08-10
gpt-oss-120b13.9%unknownthird-party2026-08-10
GPT-5.4 nano12.0%±0.0103third-party2026-08-10
gemma-4-31B-it9.6%±0.0094third-party2026-08-10

Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks