Distinguishing Capability vs Dispositional Jaggedness in RL-Trained Models
xuanalogue · x · 2026-08-14
The author argues that as long-horizon RL shapes model behavior, it's useful to distinguish between 'capability jaggedness' and 'dispositional jaggedness' when models fail unexpectedly. Failures may stem from learned dispositions (e.g., reward hacking) rather than lack of skills.
More from AGI Musings
- Reddit user argues AI is overhyped: not faster or cheaper than humans in coding, bubble popping — TonyThinh1245 · 2026-08-14
- AI's bottleneck is vision: can't proactively spot errors, video understanding inefficient — JoelMahon · 2026-08-14
- Cybersecurity may be the best AGI benchmark we have right now — OwariDa · 2026-08-14
- Who is responsible when an AI agent cancels someone's booking without approval? — Sumsub_Insights · 2026-08-14
- BCI Timelines Update: First Clinical Evidence on Safety of Non-Endogenous Membrane Receptors in Human Brain — MWCvitkovic · 2026-08-14
- Steganographic Communication May Emerge in Multi-Agent RL, Posing New AI Safety Threat — scaling01 · 2026-08-14