'Reasoning depth' estimate flags Sol 6.1 as a suspected looper — but the metric itself may be unreliable
scaling01 · x · 2026-09-30
Using a 'true reasoning depth' estimate, the author ran a spot-the-looper analysis: 5.6 luna/sol and 6 luna/sol show no looping, while 6.1 sol and 6 astra look suspicious. But he concludes the simplest explanation is that the metric isn't very good — and a single loop doesn't double reasoning depth anyway.
Related event: Benchmark sleuths suspect Grok 6.1 of answer-reusing loops(3 posts)→
More from Models
- Sam Altman asks how OpenAI should charge for Codex, $500/mo plan speculation spreads — chaumian · 2026-09-30
- ChatGPT Pro's $200 plan reportedly includes 62,500 Codex credits expiring Dec 31 — chaumian · 2026-09-30
- User calculates 62,500 credits ≈ $2,500 of GPT-6.1 API usage, calling the new plan a big cut — chaumian · 2026-09-30
- Reddit user reports surprise 62,500 credit grant, about 15x their monthly plan allowance — IronDarbe · 2026-09-30
- Users question why ChatGPT chat mode still misses the newest GPT models — koltregaskes · 2026-09-30
- User receives 62,500 credits worth ~$2,500, equal to 12.5 months of the $200 Pro plan — kimmonismus · 2026-09-30