TerminalBench-Science released: Opus 5 achieves 30% pass rate
BenBlaiszik · x · 2026-08-28
TerminalBench-Science v0.1 is released, a benchmark built by scientists to evaluate agents on scientific tasks. Initial results show Claude Opus 5 (Claude Code) leading with a 30.0% pass rate, followed by GPT-5.6 Sol (Codex) at 22.4% and Claude Fable 5 at 21.4%. GPT 5.6 variants and Opus 5 dominate the Pareto frontier for cost vs. task resolution.
More from Research
- DSPy.Flex Optimizes Python Modules with Fewer LLM Calls — lateinteraction · 2026-08-28
- MIT Framework PottsMPNN Designs Proteins Beyond Natural Structures — nordicinst · 2026-08-28
- Podcast: Discussing Python Tooling Handbook and the Value of Blogs in the Age of Agents — tdhopper · 2026-08-28
- MIT's PottsMPNN Framework Advances Protein Design Beyond Natural Sequences — MIT News AI · 2026-08-28
- Debate on the Term "Double Blind" in Model Evaluation Context — BlancheMinerva · 2026-08-28
- Debugging exact DDP checkpoint resumption: CUDA atomics, Triton autotune and bucket layout pitfalls — ReinforcedKnowledge · 2026-08-28