AI Evaluation Faces Paradigm Shift as Anti-Cheating and Costs Surge

ziv_ravid · x · 2026-07-30

As models improve, preventing them from cheating benchmarks becomes increasingly difficult and critical. Academia has been priced out of both training and evaluation, and even AI labs struggle with variance reduction due to high costs.

Furthermore, there is a philosophical shift towards Agentic benchmarks. We are now evaluating the model-harness pair, which complicates the creation of controlled environments, a tension highlighted by the recent ARC-AGI-3 and OpenAI situation.

Original post →

More from Research

Research channel →