'Reasoning depth' estimate flags Sol 6.1 as a suspected looper — but the metric itself may be unreliable

scaling01 · x · 2026-09-30

Using a 'true reasoning depth' estimate, the author ran a spot-the-looper analysis: 5.6 luna/sol and 6 luna/sol show no looping, while 6.1 sol and 6 astra look suspicious. But he concludes the simplest explanation is that the metric isn't very good — and a single loop doesn't double reasoning depth anyway.

Related event: Benchmark sleuths suspect Grok 6.1 of answer-reusing loops(3 posts)→

Original post →

More from Models

Models channel →