Terminal-Bench-Science: the best AI agents solve only 30% of real research tasks

Share
Terminal-Bench-Science: the best AI agents solve only 30% of real research tasks

A new benchmark built from researchers' own lab work is drawing the first honest line between AI agents that can talk about science and agents that can actually do it. The answer is humbling: the best system in the world clears fewer than a third of the tasks.

Terminal-Bench-Science 0.1 launches today with 70 tasks drawn from real scientific workflows across the life, physical, Earth, mathematical, and engineering sciences. Built by the Harbor framework team with Stanford's AI lab and domain experts, it extends the Terminal-Bench methodology that has driven progress in software-engineering agents into scientific research. Scientists — not model developers or data vendors — wrote and reviewed the tasks, which span data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, and scientific machine learning.

The headline number is deliberately sobering. Every model ran three independent trials on each task, and Claude Opus 5 with Claude Code came out on top by resolving just 30.0% of them. GPT-5.6 Sol with Codex followed at 22.4%, and Claude Fable 5 at 21.4%. Claude Opus 4.8 sits mid-pack at 10.5%, while GPT-5.6 Terra, Kimi K3, and Grok 4.6 all clear less than 10%. The strongest open model is GLM 5.3 at 8.1%, and GPT-5.6 Luna brings up the rear at 3.3%.

The gap between this benchmark and its sibling is the point. Compared with Terminal-Bench 3.0, resolution rates drop more than ten percentage points for every model evaluated on both — and the authors say that's deliberate: tasks were calibrated to challenge the newest frontier models, so the room for progress is real and measurable rather than already saturated. Cost matters too. GPT-5.6 Sol matches Fable 5's performance at under a third of the cost ($4.2k versus $14.2k across the full suite), while Opus 5 reaches the highest resolution at $7.0k. Only Kimi K3 and Opus 5 land on both the cost and token-efficiency Pareto frontiers.

The takeaway here isn't that AI is failing at science — it's that "useful research assistant" is still a much harder bar than "writes good code," and we finally have a yardstick that says exactly how far the frontier has to move. A benchmark that pulls real, verifiable workflows from working scientists, and refreshes as the frontier advances, is the kind of instrument the field has been missing. Expect it to become a standard citation whenever a lab claims its model is a scientist.

Would you trust an AI agent that clears less than a third of your lab's real workflows? Tell us in the comments.

Sources: Terminal-Bench-Science announcement · Terminal-Bench-Science (GitHub) · Harbor framework announcement · Snorkel AI