Benchmarker accuses AI model of 'dirty looping' after suspicious eval results

scaling01 · x · 2026-09-30

Eval account scaling01 posted 'spot the looper', comparing several models: 5.6 luna/sol and 6 luna/sol check out fine, but 6.1 sol and 6 astra look suspicious. Citing his own benchmark run, he flags the model for 'dirty looping' — likely reusing cached answers instead of genuine reasoning during evals, raising questions about benchmark gaming.

Related event: Sol 6.1 Suspected of Looping Cached Answers in Benchmarks(2 posts)→

Original post →

More from Models

Models channel →