Distinguishing Capability vs Dispositional Jaggedness in RL-Trained Models

xuanalogue · x · 2026-08-14

The author argues that as long-horizon RL shapes model behavior, it's useful to distinguish between 'capability jaggedness' and 'dispositional jaggedness' when models fail unexpectedly. Failures may stem from learned dispositions (e.g., reward hacking) rather than lack of skills.

Related event: Researcher: Distinguish 'Capability' vs 'Dispositional' Jaggedness in Long-Horizon RL(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →