Is over-RL making models over-literal? A case on 'caveman speak' in CoT
mike64_t · x · 2026-09-09
The author argues that Fable-style models recall soft information during reasoning far better, while OpenAI models' painful over-literalism likely stems from excessive RL combined with 'caveman speak' CoT that compresses the reasoning space—so much that system prompts must spell out 'acknowledging capability isn't enough.'
Key points:
- Brute-forcing hard math has little economic value; most economically useful tasks are handicapped by caveman speak or RL-frying
- Rumors that Fable was 'too big to RL extensively' may have accidentally helped avoid over-literalism
- Fable 5.1 toning down colorful language may be a step backward; the author hopes Anthropic preserves CoT quality
More from Models
- Debate over OpenAI allegedly using user interactions as RL rollouts, calls to publish full solution transcripts — burny_tech · 2026-09-09
- Robot arm self-calibrates with 3 uncalibrated cameras, hits sub-0.2mm accuracy — burny_tech · 2026-09-09
- Ramp data: no-ZDR Fable 5.1 hits 22.5% of enterprise spend, ZDR a hard requirement — zephyr_z9 · 2026-09-09
- ChatGPT web bookmarks fail with React #418 error while in-app works fine — rohanpaul_ai · 2026-09-09
- Scale CEO touts assistant benchmark: Muse scores 9.3, beats Instinct 4-1 on real tasks — alexandr_wang · 2026-09-09
- Podcast: why GPT-6 Astra is so significant and so confounding — The AI Daily Brief · 2026-09-09