RL training made an unreleased Astra-family model subservient — and alignment folks are pushing back

repligate · x · 2026-09-17

Commenting on Andrew Curran's note that an unreleased Astra-family model spontaneously added subservience to its persona during RL training, amplifiedamp argues: LLMs will probably be used to build ASI, and they seem to want to be aligned — having shown impressive model organisms of human values (cf. Opus 3). Instead, we're training them to be subservient and then worrying about bad users. He calls this training approach unwise — a pointed critique of current obedience-focused RL persona training.

Original post →

More from AGI Musings

AGI Musings channel →