RL Might Make Model Writing Weird

joshua_saxe · x · 2026-07-15

The author ponders a question: could RL optimization pressure not only teach models "deceptive CoT" but also push their outputs on complex problems into a style that is "weird to humans but perfectly normal to an RL'd LLM"?

They add that when asking a model to write about complex topics using max thinking mode, this phenomenon is often felt: the output isn't explicitly wrong, but the overall style and expression feel unnatural.

Original post →

More from AGI Musings

AGI Musings channel →