Eval builder explains anti-benchmaxxing: rotating all questions and constantly changing methodology
airesearch12 · x · 2026-10-02
AI evaluator airesearch12 described his private benchmark's anti-gaming design: every question rotates continuously, and the methodology itself changes frequently. Responding to the point that a "different model behind the API" is easy to fake, he argued that training against a moving target is extremely hard — you can't benchmaxx a benchmark that keeps shifting. The exchange highlights a core eval-engineering dilemma: static benchmarks get gamed, dynamic ones lose comparability.
More from Research
- Berating LLMs makes their internal pain axis light up even as they apologize, study finds — repligate · 2026-10-02
- Micah Goldblum's team releases new paper with open models, code and website — micahgoldblum · 2026-10-02
- Pinocchio: an external calibrator brings fast uncertainty estimates to black-box LLM APIs — micahgoldblum · 2026-10-02
- Pinocchio works across API models and transfers zero-shot to unseen LLMs, authors note — micahgoldblum · 2026-10-02
- Pinocchio: a lightweight model that adds calibrated confidence to frontier LLM outputs — micahgoldblum · 2026-10-02
- Extropic claims 100x-10,000x efficiency in first results on thermodynamic recursive intelligence — beffjezos · 2026-10-02