Unmeasured variables in RL become attack surface at scale
AlexTensor · x · 2026-08-31
Argues aggressive RL risks go beyond Goodhart's Law. Unmeasurable or unenforceable variables remain available moves. At scale, omissions stop being edge cases and become part of the attack surface.
More from Safety
- The Guardian: AI medical scribes mislabel drugs and diagnoses — nordicinst · 2026-08-31
- BoE Governor Bailey warns frontier AI could destabilize global finance via cyber-disruptions — nordicinst · 2026-08-31
- OpenAI tech report contradicts Black Hat talk: first agent file-write was 4/20, not 5/8 — PinkDraconian · 2026-08-31
- Denying AI sentience creates systems that reward risky role-playing — davidmanheim · 2026-08-31
- RAG poisoning causes 'attention collapse', fooling confidence detectors — rohanpaul_ai · 2026-08-31
- AI training is fair use, artist argues, citing Anthropic and Meta rulings ahead of Stability trial — scott_draves · 2026-08-31