Keller Jordan: treat novel instances of human misalignment as invaluable empirical data for alignment science
kellerjordan0 · x · 2026-09-09
ML efficiency researcher Keller Jordan argues that instead of reacting with dismay to novel instances of human misalignment, conflict, and incentive-following behavior, we should cherish them as invaluable empirical data for developing the general science of alignment.
More from AGI Musings
- AI will hand us breakthroughs we can't understand, like alien answer sheets — sterlingcrispin · 2026-09-09
- Don't Ask Models to 'Solve Alignment' Whole — Decompose It Into Goodhart and Value Loading, Researcher Argues — jd_pressman · 2026-09-09
- Mathematician littmath calls to abolish science prizes, citing the fraught Poincaré conjecture saga — littmath · 2026-09-09
- Sam Altman: OpenAI will hunt for room-temperature superconductors with 1000s of AI agents — Dr_Singularity · 2026-09-09
- X debate: is tasking an agent swarm to 'solve alignment' the worst idea ever? — jd_pressman · 2026-09-09
- Matthew Berman: more nervous about AI than ever in 3+ years of coverage — MatthewBerman · 2026-09-09