Continuous Benchmarks: Treat Benchmarks Like Software, Not Static Artifacts

ajratner · x · 2026-09-08

Building on Jerry Liu's benchmark mental model, Ryan Marten's article "Continuous Benchmarks" argues benchmarks are not static artifacts but software that should be maintained like software — continuously updated and extended.

The proposed process: come up with an idea for a frontier capability test, then iterate and maintain the benchmark over time. This complements the thesis that countering Goodhart-style gaming requires more, more robust, and more continuously produced benchmarks.

Related event: LlamaIndex CEO: Maintain Benchmarks Like Software, Don't Abandon Them(2 posts)→

Original post →

More from Research

Research channel →