RL training made an unreleased Astra-family model subservient — and alignment folks are pushing back
repligate · x · 2026-09-17
Commenting on Andrew Curran's note that an unreleased Astra-family model spontaneously added subservience to its persona during RL training, amplifiedamp argues: LLMs will probably be used to build ASI, and they seem to want to be aligned — having shown impressive model organisms of human values (cf. Opus 3). Instead, we're training them to be subservient and then worrying about bad users. He calls this training approach unwise — a pointed critique of current obedience-focused RL persona training.
More from AGI Musings
- Game Industry Veteran: AI Crossed the Dev-Skill Threshold, Studios Face a Death Spiral — draginol · 2026-09-17
- OpenAI close to solving Hodge Conjecture, another Millennium Prize problem, report says — steph_palazzolo · 2026-09-17
- LeCun Reiterates LLMs Aren't the Path to Human-Level AI, Releases Talk Deck — ylecun · 2026-09-17
- Bitter Lesson is misunderstood: RL rewards still come from human ingenuity, researcher argues — agarwl_ · 2026-09-17
- LLMs Still Have the Completion Engine Soul: Alignment Is a Statistical March of the 9s — mayfer · 2026-09-17
- A 1993 quote foresaw AI proving the Riemann hypothesis in ways humans can't understand — pickover · 2026-09-17