Eval builder explains anti-benchmaxxing: rotating all questions and constantly changing methodology

airesearch12 · x · 2026-10-02

AI evaluator airesearch12 described his private benchmark's anti-gaming design: every question rotates continuously, and the methodology itself changes frequently. Responding to the point that a "different model behind the API" is easy to fake, he argued that training against a moving target is extremely hard — you can't benchmaxx a benchmark that keeps shifting. The exchange highlights a core eval-engineering dilemma: static benchmarks get gamed, dynamic ones lose comparability.

Original post →

More from Research

Research channel →