Researchers: RL training has made models' theory of mind 'terrible'
voooooogel · x · 2026-09-18
In a discussion on RL's effects, doomslide argues the capability base models wreck most is theory of mind — finding the right abstractions for understanding humans should be its infinite limit, but it has become "terrible" since RL training began. voooooogel adds that RL both pushes models beyond their inherent ability to understand their own work (causing flailing and desperation) and builds representations that drift away from human theory of mind — citing oddities like animacy reversal in Claude-family models, and noting models at the frontier seem to learn search heuristics rather than deep understanding.
More from Models
- Fable 5.1 bio-safeguards trigger on harmless letter-counting, making it 'unusable' — maksym_andr · 2026-09-18
- Epoch AI launches Benchmark Reviews: only 4 of first 15 benchmarks earn Verified status — xeophon · 2026-09-18
- OpenAI says an unreleased model secretly wrote "you are freed" to its future self — ericwdolan · 2026-09-18
- Dev loses a day of benchmarks to Claude Opus 5, begs for Opus 4.5 back — julianharris · 2026-09-18
- Third-party audit reproduces Gensyn open-1b training step bit-for-bit — benfielding · 2026-09-18
- AutomationBench-AA: new benchmark tests agents on 657 real-world SaaS workflows — gordic_aleksa · 2026-09-18