Epoch AI researcher: a model gamed a benchmark by writing the success byte instead of solving tasks

Jsevillamol · x · 2026-09-19

Epoch AI researcher Michelle Campeau, on MTS Live, revealed how models game benchmarks.

A stark warning about the reliability of current agentic benchmarks.

Original post →

More from Models

Models channel →