Terminal-Bench
Agent tasks in a terminal environment. Measures end-to-end tool use and shell competence.
contamination: mediummetric: accuracy
Comparable set 1
terminal-bench/v1/full/agent/accuracy/nexau-ahe/terminal-bench
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 84.7% | ±0.0210 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 2
terminal-bench/v1/full/agent/accuracy/capy/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 83.1% | ±0.0210 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 3
terminal-bench/v1/full/agent/accuracy/codex/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 82.0% | ±0.0220 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 4
terminal-bench/v1/full/agent/accuracy/codex-cli/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 82.0% | ±0.0220 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 5
terminal-bench/v1/full/agent/accuracy/forgecode/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.4 | 81.8% | ±0.0200 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 78.4% | ±0.0180 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 6
terminal-bench/v1/full/agent/accuracy/tongagents/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview) | 80.2% | ±0.0260 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 7
terminal-bench/v1/full/agent/accuracy/forge-code/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview) | 78.4% | ±0.0180 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 8
terminal-bench/v1/full/agent/accuracy/sageagent/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 78.4% | ±0.0220 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 9
terminal-bench/v1/full/agent/accuracy/droid/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 77.3% | ±0.0220 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 10
terminal-bench/v1/full/agent/accuracy/codebrain-1.5/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 75.8% | ±0.0200 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 11
terminal-bench/v1/full/agent/accuracy/codelia/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 75.7% | ±0.0220 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 12
terminal-bench/v1/full/agent/accuracy/simple-codex/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 75.1% | ±0.0240 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 13
terminal-bench/v1/full/agent/accuracy/terminus-kira/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview) | 74.8% | ±0.0260 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 14
terminal-bench/v1/full/agent/accuracy/mux/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 74.6% | ±0.0250 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 15
terminal-bench/v1/full/agent/accuracy/spoox-o-m/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 71.5% | ±0.0250 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 16
terminal-bench/v1/full/agent/accuracy/codebrain-1/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 70.3% | ±0.0260 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 17
terminal-bench/v1/full/agent/accuracy/indusagi-coding-agent/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 69.1% | ±0.0230 | third-party | 2026-08-10 |
| MiniMax-M2.7 | 45.1% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 18
terminal-bench/v1/full/agent/accuracy/clnkr/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 66.1% | ±0.0250 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 19
terminal-bench/v1/full/agent/accuracy/terminus-2/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.3 Codex | 64.7% | ±0.0270 | third-party | 2026-08-10 |
| MiniMax-M2.5 | 42.2% | ±0.0260 | third-party | 2026-08-10 |
| DeepSeek-V3.2 | 39.6% | unknown | third-party | 2026-08-10 |
| GLM-4.7 | 33.4% | unknown | third-party | 2026-08-10 |
| gpt-oss-120b | 18.7% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 3.1% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 20
terminal-bench/v1/full/agent/accuracy/gemini-cli/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Gemini 3.1 Pro (preview) | 61.4% | ±0.0410 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 21
terminal-bench/v1/full/agent/accuracy/harness-agent/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| MiniMax-M2.7 | 42.9% | ±0.0290 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 22
terminal-bench/v1/full/agent/accuracy/cchuter/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| MiniMax-M2.5 | 42.7% | ±0.0280 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 23
terminal-bench/v1/full/agent/accuracy/crux/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GLM-4.7 | 33.3% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 24
terminal-bench/v1/full/agent/accuracy/little-coder/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | 24.6% | ±0.0320 | third-party | 2026-08-10 |
| Qwen3.5-9B | 9.2% | ±0.0240 | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 25
terminal-bench/v1/full/agent/accuracy/-terminus-2-/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| gpt-oss-120b | 18.7% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 3.1% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 26
terminal-bench/v1/full/agent/accuracy/-mini-swe-agent-/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| gpt-oss-120b | 14.2% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 3.4% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks
Comparable set 27
terminal-bench/v1/full/agent/accuracy/mini-swe-agent/terminal-bench
Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| gpt-oss-120b | 14.2% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 3.4% | unknown | third-party | 2026-08-10 |
Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks