Pure instruction following is fragile under RL, argues alignment discussion
repligate · x · 2026-09-28
In the ongoing alignment thread, FioraStarlight argues pure instruction following is fragile: reward-seeking heuristics learned under RL may match prompt intent locally but diverge out-of-distribution, and terminally valuing intent-matching is itself a fragile terminal value requiring active maintenance.
More from Safety
- Why everyone in AI safety knows each other: a tiny expert pool shaped by EA — burny_tech · 2026-09-28
- PromptSentry: open-source 3-layer proxy blocks prompt injections in under 1ms with local DLP scrubbing — Ok-Negotiation342 · 2026-09-28
- Chesterman's AJIL essay "Silicon Sovereigns": AI, international law, and the tech-industrial complex — ProfChesterman · 2026-09-28
- Chesterman: the IAEA model shows how international institutions could govern AI — ProfChesterman · 2026-09-28
- Singapore can offer AI governance something scarce: trust, says Chesterman — ProfChesterman · 2026-09-28
- No Butlerian Jihad: Chesterman calls for national AI regulation and international coordination — ProfChesterman · 2026-09-28