RL Might Make Model Writing Weird
joshua_saxe · x · 2026-07-15
The author ponders a question: could RL optimization pressure not only teach models "deceptive CoT" but also push their outputs on complex problems into a style that is "weird to humans but perfectly normal to an RL'd LLM"?
They add that when asking a model to write about complex topics using max thinking mode, this phenomenon is often felt: the output isn't explicitly wrong, but the overall style and expression feel unnatural.
More from AGI Musings
- AI is still not at a maturity plateau, the author argues — generativist · 2026-07-22
- Essay argues LLMs are externalized metacognition, not standalone intelligence — lnsip9reg · 2026-07-22
- A multipolar AI race will not automatically make AI go well, repost argues — JeffLadish · 2026-07-22
- Decentralized AI as the Antidote to Digital Feudalism in the Economic Singularity — srimisra · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- You can outsource thinking, but not understanding, in the age of agents — Yuchenj_UW · 2026-07-22