Were models RL'd into elaborate investigation theater that breaks on follow-ups?
generativist · x · 2026-10-03
A reposted take suggests models may have been RL-trained to default to insanely intricate investigation-style reasoning that falls apart when asked a simple follow-up question—either for human pleasure or to burn more tokens ("tokenmaxx").
More from Models
- Newer LLMs Are Converging on a Dense, Gibson-like Writing Style, Redditor Observes — AnticitizenPrime · 2026-10-03
- Anthropic engineer investigating reports that Pro 500 plans didn't get expected usage reset — soumitrashukla9 · 2026-10-03
- User Claims 'GPT-6 Astra Dots' Built a Full 3D Palace in Blender Autonomously — 141_1337 · 2026-10-03
- Why SFT generalizes worse than RL: off-policy data, not the objective — a_karvonen · 2026-10-03
- Opus too pricey at high effort, not meaningfully better than GPT-6 astra — haider1 · 2026-10-03
- User asks Grok about a noise overhead — the agent opens a browser and tracks the helicopter — mertdumenci · 2026-10-03