Blog
Updates and announcements from the Terminal-Bench team
Terminal-Bench 4.0
Calibrating task resources, fixing tasks, and removing saturated tasks
TERMINAL-BENCH-SCIENCE 0.1
Evaluating AI agents on scientific research workflows
Continuous Benchmarks
Benchmarks are software and should be maintained like software
Terminal-Bench 3.0
Terminal-Bench 3.0 measures agent abilities at the frontier
Harbor-Index
A lightweight, diverse, and difficult benchmark for agentic evaluation
Terminal-Bench Challenges
Long-horizon, token-intensive, single-task benchmarks
Terminal-Bench 2.1
A revision of Terminal-Bench 2.0 that fixes 28 tasks
Leaderboard Integrity Update
New policies to address cheating and reward hacking
What Makes a Good Terminal Bench Task
Building adversarial, difficult, and legible benchmark tasks
Terminal-Bench-Science Call for Contributions
Contribute your scientific workflows as benchmark tasks
Terminal-Bench 3.0 Call for Contributions
Collaborate on a new frontier of challenging computer-based tasks
Terminal-Bench 2.0 and Harbor
A harder, better Terminal-Bench + a new agent eval & optimization package
Leaderboard Integrity and Timeouts
A clarification on our leaderboard time constraints
The Terminal-Bench Dataset Registry
Evaluate agents on popular benchmarks and distribute new ones
Warp scores a new SOTA on Terminal-bench
Warp debuts at #1 on Terminal-Bench, resolving 52% of tasks
Task Spotlight: Scientific Computing and Cryptography
Highlighting two new tasks for Terminal-Bench
Terminal-Bench on the Claude 4 Model Card
Anthropic features Terminal-Bench and sets a new SOTA
Terminal-Bench
An evaluation framework and benchmark for agents in the terminal
Terminus
A research-preview agent for evaluating language models in the terminal