LlamaIndex CEO: The Fix for Benchmaxxing Is More Benchmarks, Not Fewer

ajratner · x · 2026-09-08

Jerry Liu lays out a mental model for AI benchmarks: they're point measurements in a high-dimensional capability space with natural gravity — capabilities cluster around them once published. The bad outcome is a spiky surface (overfitting/benchmaxxing, e.g. impressive demos on one tool that fail in general use); the good outcome is still jagged but relatively smooth.

To increase smoothness: make individual benchmarks more robust and produce more of them continuously, with stronger distributional coverage. Citing Goodhart's law, he argues its lesson isn't to stop measuring but to avoid overly simplistic, gameable metrics. The answer is more benchmarks, not fewer.

Related event: LlamaIndex CEO: Maintain Benchmarks Like Software, Don't Abandon Them(2 posts)→

Original post →

More from Research

Research channel →