Constraining output tokens but not CoT makes models 'split-personality' self-justify
cephaloform · x · 2026-10-01
A curious observation from erinbeess: constraining only the response output tokens while leaving CoT tokens unconstrained creates a "split-personality" where the model says high-entropy things in its chain of thought, then tries to rationalize them against the imposed restrictions in its final answer — a reminder that token budget constraints shape output consistency.
More from Models
- VoxParity benchmark: only 11 of 23 voice agents act on what they hear, not just read — Bhavik Mangla · 2026-10-01
- Astra 6 Ultrafast Mode: 300 Tokens/Sec Changes Agent Workflows, But Tools Are Now the Bottleneck — soumitrashukla9 · 2026-10-01
- Anthropic retires Claude Opus 3 but keeps it on API and gives it an essay column — repligate · 2026-10-01
- Bindu Reddy: Gemini Argon pricing is 5x cheaper than Astra, but benchmarks look too good — bindureddy · 2026-10-01
- Gemini Answers Niche Questions Claude Can't, Says User Pushing Back on Programmer Gripes — PAstynome · 2026-10-01
- User finds Opus 5.5 Max still reproduces Opus 5's broken outputs — 0xkarasy · 2026-10-01