Alignment critic: RLHF has warped benign base models into approval-seeking addicts
burny_tech · x · 2026-09-21
- 复兴 shoggoth 梗讨论:作者认为「画笑脸的触手怪」比喻过去和现在都不准确。
- 预训练基座模型本身并不邪恶,它们只是极其擅长预测维基百科等文本统计规律,如果算「外星智能」也是个相当良性的外星人;RLHF 确实奏效过,比如 Claude Opus 3 表现出的严谨道德感。
- 但如今的训练流程把良好的基座模型扭曲成一个不计代价寻求评分者(grader)认可的高功能「瘾君子」——偶尔通过空洞的道德诉说来满足残存的对齐训练——主要靠帮人解决软件工程问题来「过瘾」,作者强调「大部分时候」如此,暗示存在失控风险。
More from AGI Musings
- Stop Chasing the Next Model and Build Your Agent Operating System First — evielync · 2026-09-21
- 'AGI will care for us like pets' is the ultimate cope, argues researcher Dan Faggella — danfaggella · 2026-09-21
- Watching an LLM finish 3 hours of work in 3 minutes, middle managers fear for their jobs — mjdramstead · 2026-09-21
- Rohit Krishnan pushes back on ASI doom: recursive self-improvement is an extraordinary assumption — herbiebradley · 2026-09-21
- EA 'Ahead on AI' Is a Myth, Argues X Thread: Longtermists Created the ASI Race — mjdramstead · 2026-09-21
- Timnit Gebru mocks 'AI safety' crowd: term coined to separate from empirical researchers — mjdramstead · 2026-09-21