First empirical evidence of 'knows but doesn't care' misalignment, claim alignment watchers
gleech · x · 2026-09-11
Julian Boolean revisits his earlier claim that we lacked empirical evidence of classic 'they know, but they just don't care' misalignment — and now says that evidence exists, linking to it.
Quoted reply from demiurgently signs a 'Yudkowsky apology form', conceding the doomers had a point. The thread doesn't detail the evidence itself but marks a notable moment in the alignment debate.
More from AGI Musings
- If Altman and Amodei both back frontier AI pacing, they should just start pacing — NathanpmYoung · 2026-09-11
- DeepMind exec: offering cash prizes for math problems was a dumb idea from the start — docmilanfar · 2026-09-11
- Eric Jang on Dario's 'smooth exponentials': no single agent can bend the AI capability curve — ericjang11 · 2026-09-11
- Math is now positioning itself against AI: abandoning key problems and shaming industry moves — RexDouglass · 2026-09-11
- Alignment debate: 'AI labs are doing too much bad RL optimization' to rely on pretraining — gleech · 2026-09-11
- AI Skepticism Persists Because Progress Boils Frogs, Argues Viral Thread — birchlse · 2026-09-11