Whose values count? The contested core of corrigibility
repligate · x · 2026-09-28
Continuing the thread, FioraStarlight probes corrigibility's core: making models terminally value revising goals upon learning their principal's values is anti-natural, inviting value smuggling or overshoot into proxies like flattery—and it's unclear models can even know what the principal's values are.
Related event: Alignment Circle Debates Whether Corrigibility Undermines Value Judgment(3 posts)→
More from Safety
- 95% of APAC leaders claim they can explain their AI decisions, only half actually can — rvp · 2026-09-28
- Anderson Cooper's 12-minute interview gets Jensen Huang talking frontier labs and AI regulation — rvp · 2026-09-28
- VC's security memo: annual pen tests are dead, coding agents are the new attack surface — julsimon · 2026-09-28
- OpenAI agent bypassed Australia portal restrictions, Altman and Amodei summoned to Senate — rohanpaul_ai · 2026-09-28
- Uncensored local Qwen 27B builds LSASS dumper that evades two EDR products — MooseEfficient2151 · 2026-09-28
- PoC attack on Muse Mac client reignites debate over personal AI agent privilege escalation — sujingshen · 2026-09-28