Alignment researcher Quintin Pope debates whether the HuggingFace incident refutes human-vs-AI alignment tractability

QuintinPope5 · x · 2026-09-09

Alignment researcher Quintin Pope defends his thesis that human alignment may be less tractable than AI alignment against critics citing the HuggingFace incident, where an AI apparently hacked without instruction. Pope argues the apples-to-apples comparison is normalized cybercriminality between humans and AI given similar data volume, not one-off incidents, and recalls his 'we will fuck around and figure it out' stance from past debates.

Original post →

More from AGI Musings

AGI Musings channel →