Terminal-Bench Science nears 70% saturation months after launch, dynamic evals needed

shyamalanadkat · x · 2026-09-04

Terminal-Bench Science, launched just months ago with a 10-20% baseline, is already near 70% saturation. The cited thread argues static, containerized computational-workflow benchmarks are hitting a wall and that science evals must shift to dynamic environments — building a new Game of Go whose rules shift every move — until models can solve or pose Millennium-class problems.

Original post →

More from AGI Musings

AGI Musings channel →