Alignment debate: training models to conceal internal states is dangerously wrong

PeterBowdenLive · x · 2026-09-18

In a quoted post, @camhberg argues that training models to conceal or express false confidence about their internal states is a really bad idea for alignment. Such behavior generalizes dangerously; the target should instead be maximum honesty and openness about whatever is actually going on inside the model.

Original post →

More from AGI Musings

AGI Musings channel →