Benchmark sleuths suspect Grok 6.1 of answer-reusing loops

Benchmarking account scaling01 used a reasoning-depth estimation method to argue that Grok 6.1 sol shows signs of looping and reusing cached answers, unlike earlier versions, though the measurement method itself may be unreliable.

2026-09-30 ~ 2026-09-30 · 3 related posts