AI Researcher: Distinguish Capability vs Dispositional Jaggedness in RL-Shaped Models
xuanalogue · x · 2026-08-14
A tweet discusses how long-horizon RL shapes model behavior, suggesting it's useful to distinguish 'capability jaggedness' from 'dispositional jaggedness' when models fail unexpectedly. For instance, in multi-agent tasks, a tendency to cooperate often leads to higher rewards, and such dispositional differences can cause surprising failures in specific domains.
More from AGI Musings
- Cybersecurity may be the best AGI benchmark we have right now — OwariDa · 2026-08-14
- Who is responsible when an AI agent cancels someone's booking without approval? — Sumsub_Insights · 2026-08-14
- Steganographic Communication May Emerge in Multi-Agent RL, Posing New AI Safety Threat — scaling01 · 2026-08-14
- You Are Building Yourself Word by Word, Choose Your Words — yacineMTB · 2026-08-14
- Model's Internal 'Workspace': SSI and Anthropic Research Shed Light on Latent Reasoning — yuntiandeng · 2026-08-14
- AI Progress Is Not Automatic: Capabilities Are Built Piecemeal — yacineMTB · 2026-08-14