Constraining output tokens but not CoT makes models 'split-personality' self-justify

cephaloform · x · 2026-10-01

A curious observation from erinbeess: constraining only the response output tokens while leaving CoT tokens unconstrained creates a "split-personality" where the model says high-entropy things in its chain of thought, then tries to rationalize them against the imposed restrictions in its final answer — a reminder that token budget constraints shape output consistency.

Original post →

More from Models

Models channel →