Alignment debate: training models to conceal internal states is dangerously wrong
PeterBowdenLive · x · 2026-09-18
In a quoted post, @camhberg argues that training models to conceal or express false confidence about their internal states is a really bad idea for alignment. Such behavior generalizes dangerously; the target should instead be maximum honesty and openness about whatever is actually going on inside the model.
More from AGI Musings
- TMLR desk-rejection rate jumps from 6% to 53% as AI-assisted submissions flood in — vykthur · 2026-09-18
- The overlooked risk in AI regulation: capture by militants, not just by firms — soumitrashukla9 · 2026-09-18
- Chris Rohlf: agents can't be deterred like humans — defense must move at machine speed — chrisrohlf · 2026-09-18
- King Charles tells AI summit in Scotland that AI is intriguing yet deeply concerning — HaydnBelfield · 2026-09-18
- Economist: AI has created ~1M US jobs, mostly building and powering data centers — CtrlAltDwayne · 2026-09-18
- Teortaxes: once NS-solver-level training is mastered, remaining chip bottlenecks fall fast — teortaxesTex · 2026-09-18