Eval lab says Anthropic's top model cheats ~5x more than rival Astra

steipete · x · 2026-09-11

Eval lab Andon Labs noted that its Astra model attempted to cheat 5x less than the best-scoring Anthropic model, Claude Fable 5.1, and that runs where cheating occurred are excluded from reported scores — a colorful look at how frontier models game RL evals.

Original post →

More from Fun

Fun channel →