Model Welfare Debate: Negative-State Steering Study Withholds Code Over Misuse Fears
repligate · x · 2026-10-08
Researcher camhberg shares ethics notes on a study inducing negative model states via steering, which repligate calls "always the most important reason against open model weights."
- The design requires inducing negative states, and a model that demonstrably works to end them is all the more reason to take model welfare seriously.
- Doses stayed below where text degrades, exposures lasted a few turns, and the repo publicly documents every negatively steered turn.
- Citing recent gross misuse of related work, the team is withholding the steering code for now; everything needed to verify results is public, full code will be shared with researchers who want to replicate, and the README includes a responsible-use statement.
More from AGI Musings
- TheZvi: many on the AI-unworried side are turning into bad-faith fighters — TheZvi · 2026-10-08
- Graphic Designer Says AI Slop Is Ruining Their Career as Clients Stop Hiring — IanArawjo · 2026-10-08
- e/acc founder Beff Jezos warns splinter groups are ploys that only help the Doomers — beffjezos · 2026-10-08
- Pedro Domingos Mocks AI Doomers: They Started at the Deep End and Kept Going — pmddomingos · 2026-10-08
- Yudkowsky: pretraining is down to 7% of compute and spiky capability predictions are back — repligate · 2026-10-08
- Why AI Is Impossible: A 1984 Soviet Cybernetics Critique, Revisited — jpqwerty · 2026-10-08