Terminal-Bench

Agent tasks in a terminal environment. Measures end-to-end tool use and shell competence.

contamination: mediummetric: accuracy

Comparable set 1

terminal-bench/v1/full/agent/accuracy/nexau-ahe/terminal-bench

ModelScoreUncertaintyRun byRetrieved
GPT-5.584.7%±0.0210third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 2

terminal-bench/v1/full/agent/accuracy/capy/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.583.1%±0.0210third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 3

terminal-bench/v1/full/agent/accuracy/codex/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.582.0%±0.0220third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 4

terminal-bench/v1/full/agent/accuracy/codex-cli/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.582.0%±0.0220third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 5

terminal-bench/v1/full/agent/accuracy/forgecode/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.481.8%±0.0200third-party2026-08-10
Gemini 3.1 Pro (preview)78.4%±0.0180third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 6

terminal-bench/v1/full/agent/accuracy/tongagents/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
Gemini 3.1 Pro (preview)80.2%±0.0260third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 7

terminal-bench/v1/full/agent/accuracy/forge-code/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
Gemini 3.1 Pro (preview)78.4%±0.0180third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 8

terminal-bench/v1/full/agent/accuracy/sageagent/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex78.4%±0.0220third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 9

terminal-bench/v1/full/agent/accuracy/droid/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex77.3%±0.0220third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 10

terminal-bench/v1/full/agent/accuracy/codebrain-1.5/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex75.8%±0.0200third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 11

terminal-bench/v1/full/agent/accuracy/codelia/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex75.7%±0.0220third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 12

terminal-bench/v1/full/agent/accuracy/simple-codex/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex75.1%±0.0240third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 13

terminal-bench/v1/full/agent/accuracy/terminus-kira/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
Gemini 3.1 Pro (preview)74.8%±0.0260third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 14

terminal-bench/v1/full/agent/accuracy/mux/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex74.6%±0.0250third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 15

terminal-bench/v1/full/agent/accuracy/spoox-o-m/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex71.5%±0.0250third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 16

terminal-bench/v1/full/agent/accuracy/codebrain-1/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex70.3%±0.0260third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 17

terminal-bench/v1/full/agent/accuracy/indusagi-coding-agent/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex69.1%±0.0230third-party2026-08-10
MiniMax-M2.745.1%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 18

terminal-bench/v1/full/agent/accuracy/clnkr/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.566.1%±0.0250third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 19

terminal-bench/v1/full/agent/accuracy/terminus-2/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GPT-5.3 Codex64.7%±0.0270third-party2026-08-10
MiniMax-M2.542.2%±0.0260third-party2026-08-10
DeepSeek-V3.239.6%unknownthird-party2026-08-10
GLM-4.733.4%unknownthird-party2026-08-10
gpt-oss-120b18.7%unknownthird-party2026-08-10
gpt-oss-20b3.1%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 20

terminal-bench/v1/full/agent/accuracy/gemini-cli/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
Gemini 3.1 Pro (preview)61.4%±0.0410third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 21

terminal-bench/v1/full/agent/accuracy/harness-agent/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
MiniMax-M2.742.9%±0.0290third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 22

terminal-bench/v1/full/agent/accuracy/cchuter/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
MiniMax-M2.542.7%±0.0280third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 23

terminal-bench/v1/full/agent/accuracy/crux/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
GLM-4.733.3%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 24

terminal-bench/v1/full/agent/accuracy/little-coder/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
Qwen3.6-35B-A3B24.6%±0.0320third-party2026-08-10
Qwen3.5-9B9.2%±0.0240third-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 25

terminal-bench/v1/full/agent/accuracy/-terminus-2-/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
gpt-oss-120b18.7%unknownthird-party2026-08-10
gpt-oss-20b3.1%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 26

terminal-bench/v1/full/agent/accuracy/-mini-swe-agent-/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
gpt-oss-120b14.2%unknownthird-party2026-08-10
gpt-oss-20b3.4%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks

Comparable set 27

terminal-bench/v1/full/agent/accuracy/mini-swe-agent/terminal-bench

Separated from the set above because the comparability key differs (version, split, shot, scoring method, or harness family).

ModelScoreUncertaintyRun byRetrieved
gpt-oss-120b14.2%unknownthird-party2026-08-10
gpt-oss-20b3.4%unknownthird-party2026-08-10

Terminal-Bench (Apache-2.0), via Epoch AI Benchmarking Hub. · license Apache-2.0 · epoch-ai · https://epoch.ai/benchmarks