Benchmark sleuths suspect Grok 6.1 of answer-reusing loops
Benchmarking account scaling01 used a reasoning-depth estimation method to argue that Grok 6.1 sol shows signs of looping and reusing cached answers, unlike earlier versions, though the measurement method itself may be unreliable.
2026-09-30 ~ 2026-09-30 · 3 related posts
- Benchmarker accuses AI model of 'dirty looping' after suspicious eval results — scaling01 · 2026-09-30
- 'Reasoning depth' estimate flags Sol 6.1 as a suspected looper — but the metric itself may be unreliable — scaling01 · 2026-09-30
- Reasoning-depth estimate casts doubt on Grok 6.1 looping gains — scaling01 · 2026-09-30