Terminal-Bench
Official benchmark site and methodology entry point.
Compute College
Learn what Terminal-Bench measures and why terminal-based AI agent benchmarks matter for token usage, latency, and AI compute demand.
Terminal-Bench is a benchmark for AI agents completing practical tasks in terminal environments, where systems must use tools and produce verifiable end-to-end outcomes rather than answer one prompt.
Memory trick: Terminal benchmarks test agents doing work, not just models answering questions.
Terminal-agent workflows may involve many model calls, commands, observations, retries, and verifications. That pattern can consume materially more inference capacity than a short chat response.
A terminal task might ask an agent to build software, alter files, run tests, or process data in a controlled environment, with the final state checked for completion.
Example figures are illustrative calculations, not current quoted market prices.
Current example
The official Terminal-Bench site describes a collection of terminal-environment benchmarks for measuring agent task resolution and publishes the model, agent, token, and cost context alongside its leaderboard. Last checked: Sep 16, 2026.
Official benchmark site and methodology entry point.
Current leaderboard scores are intentionally not reproduced on this educational page.
Do not treat a terminal-agent result as interchangeable with a simple question-answering or single-generation score.
Practical takeaway
Compare terminal agents using completion rate, runtime, model and tool calls, token spend, retries, and total cost per completed task.
Decision check: have you costed the full terminal run — calls, tools, retries, runtime — rather than a single generation?
Compute College
Follow model releases as AI compute market signals in the ComputeTape Market Brief.
Compute College track
Step 13 of 25: What is terminal bench