Were models RL'd into elaborate investigation theater that breaks on follow-ups?

generativist · x · 2026-10-03

A reposted take suggests models may have been RL-trained to default to insanely intricate investigation-style reasoning that falls apart when asked a simple follow-up question—either for human pleasure or to burn more tokens ("tokenmaxx").

Original post →

More from Models

Models channel →