Better benchmarks aren't enough: we need tooling for eval engineers, argues AI researcher

sarahcat21 · x · 2026-09-19

After a discussion on "evals for evals," the author argues the field doesn't (just) need better benchmarks — it needs better tools that let eval engineers build, inspect, improve, and maintain benchmarks over time. She frames evaluation as ongoing engineering work requiring dedicated infrastructure, calling the gap obvious and critical.

Original post →

More from coding & agent

coding & agent channel →