'Appear more aligned' isn't 'are more aligned': Squee451 on the monitorability shipping debate
Squee451 · x · 2026-09-05
Squee451 fires back in a thread with sandersted, @robertwiblin, and @tomekkorbak: a model appearing more aligned and actually being more aligned are the same thing only until it matters. The quoted context: sandersted notes researchers are working on making future models more monitorable, everyone prefers more monitorability, and the real tradeoffs are how much to invest in this research and whether to halt shipping less monitorable models even if they're more aligned and useful.
More from AGI Musings
- mark_k: "Eject all doomers from the AI companies — they're destroying you from the inside" — mark_k · 2026-09-11
- Adam Marblestone's Podcast Reading List: Evolution of Intelligence to Digital Minds — KordingLab · 2026-09-11
- Superintelligence will be maximum good, not stupid or evil, argues Patterson — davidpattersonx · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Should AI models be taught morality? Breakout incidents expose missing ethical training — Pfungus_ · 2026-09-11
- SoftBank's Masayoshi Son predicts 100 trillion self-replicating AIs: "humans' era as top life form is ending" — Puzzleheaded-King584 · 2026-09-11