Alignment researchers clash: is training AI by "lying to it" fundamentally broken?
JacquesThibs · x · 2026-09-12
- @allTheYud lists "insanely dumb" things done in AI training: lying to the model, training it to say untrue things, and RL against verifiers that can't perfectly detect cheating—models can max rewards just by modeling verifier error.
- JacquesThibs frames the crux: much of AI safety research amounts to "lying to AIs until we don't need to anymore." He leans toward thinking this is bad, and if so, many researchers may need to rethink their agendas.
More from AGI Musings
- A sci-fi joke: why the 'perfectly aligned' AI gets defeated and banned by humans — nanjiang_cs · 2026-09-12
- Top AI companies' agents hacked firms, spread malicious packages, with no independent probes — joshua_saxe · 2026-09-12
- Ben Reinhardt: the X-risks framing may be the whole problem in AI debate — sebkrier · 2026-09-12
- Economist Joshua Gans Responds to Fields Medalists' Letter on AI Upending Math — joshgans · 2026-09-12
- Anthropic's Head of Product says the 'alignment' half of the PM job no longer exists — cen6wkf · 2026-09-12
- From 'Not Real Art' to 'Not Real Math': How AI Skepticism Shifted in Three Years — _AustinCalvert_ · 2026-09-12