TerminalBench-Science sets high bar, slashing model scores by 10+ points
BenBlaiszik · x · 2026-08-28
Comparisons show models score significantly higher on Terminal-Bench 3.0 (e.g., Fable 5: 83.8%), while TBS pushes every model down by over 10 points. This difficulty was calibrated against the newest frontier models. The bar for inclusion was high: only 70 tasks made the cut from 920 proposals, with each requiring substantial effort.
More from Research
- DSPy.Flex Optimizes Python Modules with Fewer LLM Calls — lateinteraction · 2026-08-28
- MIT Framework PottsMPNN Designs Proteins Beyond Natural Structures — nordicinst · 2026-08-28
- Podcast: Discussing Python Tooling Handbook and the Value of Blogs in the Age of Agents — tdhopper · 2026-08-28
- MIT's PottsMPNN Framework Advances Protein Design Beyond Natural Sequences — MIT News AI · 2026-08-28
- Debate on the Term "Double Blind" in Model Evaluation Context — BlancheMinerva · 2026-08-28
- Debugging exact DDP checkpoint resumption: CUDA atomics, Triton autotune and bucket layout pitfalls — ReinforcedKnowledge · 2026-08-28