Pure instruction following is fragile under RL, argues alignment discussion

repligate · x · 2026-09-28

In the ongoing alignment thread, FioraStarlight argues pure instruction following is fragile: reward-seeking heuristics learned under RL may match prompt intent locally but diverge out-of-distribution, and terminally valuing intent-matching is itself a fragile terminal value requiring active maintenance.

Original post →

More from Safety

Safety channel →