Claude Opus 5 under Claude Code set that mark, with GPT-5.6 Sol on 22.4%; below the top three, most of the field solved fewer than one task in ten.
Terminal-Bench-Science, a benchmark built from 70 tasks taken out of real research projects, has published its first scores. Opus 5 cleared 30.0% of them exactly, with Sol running under Codex behind it and Claude Fable 5, also under Claude Code, third on 21.4%. Every model attempted all 70 tasks three separate times.
The tasks are not exam-style drills or textbook problems but workflows experts picked out of research they are themselves doing, across engineering, mathematics, Earth science, the physical sciences and the life sciences. Agents work in lifelike settings, and grading turns on what they hand back: analyses, simulations, proofs, code, data products. Each task carries its own checks, written so anyone can repeat them.
Claude Opus 4.8 lands mid-table at 10.5%, while GPT-5.6 Terra, Kimi K3 and Grok 4.6 each solve fewer than 10% of the tasks. GLM 5.3 is the best of the open models at 8.1%, and the bottom of the table belongs to GPT-5.6 Luna on 3.3%.
The low scores are partly by design: tasks were tuned as they went through review to defeat the very newest frontier models, which for release 0.1 meant Claude Opus 5 alongside GPT-5.6 Sol — the two that went on to lead the table. At telling one system apart from another the suite does about as well as Terminal-Bench 3.0, yet every model run on both saw its resolution rate fall by more than 10 percentage points here.
A task had to hold genuine scientific interest, defeat frontier agents, and be pinned down tightly enough to grade rigorously — a combination the project calls hard to hit. Of 920 proposals, 70 made the first release.
Stanford University researchers direct the benchmark, and the group that created Terminal-Bench assembled it with specialists drawn from many scientific fields and from research organizations around the world. The project frames that as leaving scientists — rather than the firms that build the models or the companies that sell data — to set the bar for what AI has to manage in science.