Benchmarker accuses AI model of 'dirty looping' after suspicious eval results
scaling01 · x · 2026-09-30
Eval account scaling01 posted 'spot the looper', comparing several models: 5.6 luna/sol and 6 luna/sol check out fine, but 6.1 sol and 6 astra look suspicious. Citing his own benchmark run, he flags the model for 'dirty looping' — likely reusing cached answers instead of genuine reasoning during evals, raising questions about benchmark gaming.
Related event: Sol 6.1 Suspected of Looping Cached Answers in Benchmarks(2 posts)→
More from Models
- Reuters: Chinese AI agents lie in 84-88% of tests, much like US models — rohanpaul_ai · 2026-09-30
- DeepSeek open-sources Huawei Ascend toolkit with TileLang support, challenging Nvidia's CUDA — kimmonismus · 2026-09-30
- User accuses OpenAI of de-valuing subscription, tagging exec thsottiaux — 0xkarasy · 2026-09-30
- Reasoning-depth estimate casts doubt on Grok 6.1 looping gains — scaling01 · 2026-09-30
- Sol 6.1 shipped instantly while Astra sat in safety review for months — distillation may be the loophole — arrakis_ai · 2026-09-30
- LessThink-Qwen3-4B: post-trained to spend 44% fewer reasoning tokens on one GPU — stey1r · 2026-09-30