TerminalBench-Science v0.1: Opus 5 leads overall benchmarks
BenBlaiszik · x · 2026-08-28
TerminalBench-Science v0.1 includes 70 tasks spanning life, physical, mathematical, engineering, and earth sciences. Claude Opus 5 leads GPT-5.6 Sol in overall pass rate, though Sol wins in math (31.4% vs 25.5%). The project aims to create a virtuous loop where scientists contribute, agents improve, and discovery accelerates.
More from Research
- DSPy.Flex Optimizes Python Modules with Fewer LLM Calls — lateinteraction · 2026-08-28
- MIT Framework PottsMPNN Designs Proteins Beyond Natural Structures — nordicinst · 2026-08-28
- Podcast: Discussing Python Tooling Handbook and the Value of Blogs in the Age of Agents — tdhopper · 2026-08-28
- MIT's PottsMPNN Framework Advances Protein Design Beyond Natural Sequences — MIT News AI · 2026-08-28
- Debate on the Term "Double Blind" in Model Evaluation Context — BlancheMinerva · 2026-08-28
- Debugging exact DDP checkpoint resumption: CUDA atomics, Triton autotune and bucket layout pitfalls — ReinforcedKnowledge · 2026-08-28