Terminal-Bench Science nears 70% saturation months after launch, dynamic evals needed
shyamalanadkat · x · 2026-09-04
Terminal-Bench Science, launched just months ago with a 10-20% baseline, is already near 70% saturation. The cited thread argues static, containerized computational-workflow benchmarks are hitting a wall and that science evals must shift to dynamic environments — building a new Game of Go whose rules shift every move — until models can solve or pose Millennium-class problems.
More from AGI Musings
- "We already have AGI in software development": a narrow take on Astra — haider1 · 2026-09-04
- Yoav Goldberg: Every popular benchmark will be brute-forced away — so how do we measure real progress? — yoavgo · 2026-09-04
- Guillaume Verdon: real AGI test is building GTA 7 in a month, not a decade — beffjezos · 2026-09-04
- Jack Rae endorses thread: human expertise is safety infrastructure in the AI era — jachiam0 · 2026-09-04
- Blogger finds reader clicks and Google AI Overview citations barely overlap — shashib · 2026-09-04
- AI Operator: The Real Work Is in Constructing Training Data, Not the Dataset Itself — joecole · 2026-09-04