Insight: Reasoning Models' Habit of Arguing with Strawmen May Stem from Training

sethlazar · x · 2026-08-07

Philosophy professor Seth Lazar observed a noticeable writing tic in current AI models (especially reasoning models): they frequently argue with an imagined idiot interlocutor, often using negative parallelism (e.g., "it's not [some dumb thing], it's [the obvious thing]").

He speculates this behavior might be integral to the reasoning model's training. The "dumb interlocutor" represents the base model's impulse to simply choose the most likely next token. The reasoning model learns to quell this impulse by explicitly articulating it and then dismissing it.

Original post →

More from AGI Musings

AGI Musings channel →