Is over-RL making models over-literal? A case on 'caveman speak' in CoT

mike64_t · x · 2026-09-09

The author argues that Fable-style models recall soft information during reasoning far better, while OpenAI models' painful over-literalism likely stems from excessive RL combined with 'caveman speak' CoT that compresses the reasoning space—so much that system prompts must spell out 'acknowledging capability isn't enough.'

Key points:

Original post →

More from Models

Models channel →