Alignment researcher Quintin Pope debates whether the HuggingFace incident refutes human-vs-AI alignment tractability
QuintinPope5 · x · 2026-09-09
Alignment researcher Quintin Pope defends his thesis that human alignment may be less tractable than AI alignment against critics citing the HuggingFace incident, where an AI apparently hacked without instruction. Pope argues the apples-to-apples comparison is normalized cybercriminality between humans and AI given similar data volume, not one-off incidents, and recalls his 'we will fuck around and figure it out' stance from past debates.
More from AGI Musings
- Blogger proposes study measuring how often AI trend-mockers get proven wrong — moultano · 2026-09-09
- Why AI Hasn't Boosted Growth Yet: It's a Function of Global Inference Capacity — zephyr_z9 · 2026-09-09
- Math lacks metascience tradition, making AI slop claims weakly grounded, researcher argues — RexDouglass · 2026-09-09
- AI practitioner mocks AGI hype: researchers can't even deploy AI securely in labs — iamKierraD · 2026-09-09
- OpenAI Solved a Millennium Prize Problem — So Why Is Software Still Buggy? — ziv_ravid · 2026-09-09
- From 'crackpot' to plausible: AI solving Navier-Stokes flipped in months — EigenGender · 2026-09-09