GPQA Diamond
Graduate-Level Google-Proof Q&A (Diamond split). Hard STEM questions designed to be difficult for non-experts with web search.
contamination: lowmetric: accuracy
Comparable set
gpqa-diamond/v1/diamond/0shot/exact-match/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.4 | 94.6% | ±0.0160 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 94.4% | ±0.0163 | third-party | 2026-08-10 |
| Gemini 3.6 Flash | 94.1% | ±0.0140 | third-party | 2026-08-10 |
| GPT-5.5 | 94.0% | ±0.0155 | third-party | 2026-08-10 |
| Claude Opus 5 | 93.9% | ±0.0148 | third-party | 2026-08-10 |
| GPT-5.6 Sol | 93.5% | ±0.0157 | third-party | 2026-08-10 |
| Grok 4.5 | 93.4% | ±0.0143 | third-party | 2026-08-10 |
| GPT-5.6 Terra | 93.3% | ±0.0154 | third-party | 2026-08-10 |
| Kimi-K3 | 93.1% | unknown | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 92.8% | ±0.0164 | third-party | 2026-08-10 |
| Qwen3.8-Max | 92.7% | ±0.0169 | third-party | 2026-08-10 |
| GLM-5.2 | 91.9% | ±0.0161 | third-party | 2026-08-10 |
| GPT-5.6 Luna | 91.6% | ±0.0173 | third-party | 2026-08-10 |
| Claude Opus 4.8 | 91.0% | ±0.0192 | third-party | 2026-08-10 |
| DeepSeek-V4-Flash-0731 | 91.0% | unknown | third-party | 2026-08-10 |
| Qwen3.7-Max | 90.9% | ±0.0205 | third-party | 2026-08-10 |
| DeepSeek-V4-Pro | 90.9% | ±0.0205 | third-party | 2026-08-10 |
| Kimi-K2.6 | 90.8% | ±0.0172 | third-party | 2026-08-10 |
| Claude Sonnet 5 | 90.5% | ±0.0178 | third-party | 2026-08-10 |
| Grok 4.3 | 88.8% | ±0.0196 | third-party | 2026-08-10 |
| Kimi-K2.7-Code | 87.9% | ±0.0233 | third-party | 2026-08-10 |
| GPT-5.4 mini | 86.9% | ±0.0241 | third-party | 2026-08-10 |
| Qwen3.5-397B-A17B | 86.4% | ±0.0245 | third-party | 2026-08-10 |
| Qwen3.6-27B | 85.9% | ±0.0248 | third-party | 2026-08-10 |
| Claude Fable 5 | 85.9% | ±0.0248 | third-party | 2026-08-10 |
| GLM-5.1 | 85.5% | ±0.0207 | third-party | 2026-08-10 |
| Qwen3.6-35B-A3B | 84.9% | ±0.0255 | third-party | 2026-08-10 |
| Gemini 3.5 Flash Lite | 83.3% | ±0.0266 | third-party | 2026-08-10 |
| GLM-4.7 | 83.3% | unknown | third-party | 2026-08-10 |
| Qwen3-235B-A22B-Instruct-2507 | 80.0% | ±0.0260 | third-party | 2026-08-10 |
| GPT-5.4 nano | 78.5% | ±0.0244 | third-party | 2026-08-10 |
| DeepSeek-R1-0528 | 76.3% | unknown | third-party | 2026-08-10 |
| gemma-4-31B-it | 75.8% | ±0.0305 | third-party | 2026-08-10 |
| gpt-oss-120b | 75.8% | unknown | third-party | 2026-08-10 |
| Llama-4-Maverick-17B-128E-Instruct | 67.0% | unknown | third-party | 2026-08-10 |
| phi-4 | 56.1% | unknown | third-party | 2026-08-10 |
| Llama-4-Scout-17B-16E-Instruct | 51.8% | unknown | third-party | 2026-08-10 |
| Qwen2.5-72B-Instruct | 49.1% | unknown | third-party | 2026-08-10 |
| Llama-3.3-70B-Instruct | 47.4% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 46.0% | unknown | third-party | 2026-08-10 |
| Llama-3.1-70B-Instruct | 44.2% | unknown | third-party | 2026-08-10 |
| gemma-3-27b-it | 43.9% | ±0.0354 | third-party | 2026-08-10 |
| Mistral-Nemo-Instruct-2407 | 29.9% | ±0.0210 | third-party | 2026-08-10 |
| gemma-2-9b-it | 27.5% | unknown | third-party | 2026-08-10 |
| Llama-3.1-8B-Instruct | 25.9% | unknown | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks