Quintin Pope: training's robustness to data ordering explains why backdoor behaviors survive
QuintinPope5 · x · 2026-09-29
Continuing his thread, Quintin Pope argues that given how robust LLM training seems to data ordering, one must accept that trained behaviors can survive further training with inexact prefix overlap — since "train bad on prefix x, then good otherwise" is not that different from training both simultaneously.
More from AGI Musings
- Blogger reframes the singularity: not machines passing humans, but humans surrendering moral judgment — AryHHAry · 2026-09-29
- Agent ran 89 experiments to improve a small model — 92% of gains came by experiment 44 — ccerrato147 · 2026-09-29
- AI slop papers with random math are 'roleplaying science' — and may ironically make real papers easier to publish — zouharvi · 2026-09-29
- A perfectly aligned AI would never listen to humans, argues one poster — djcows · 2026-09-29
- Garrison Lovely on Why Treating AI Labor Automation as Inevitable Poses Severe Risks — The Cognitive Revolution · 2026-09-29
- Garrison Lovely on 'Obsolete': Why Racing to Replace All Human Labor Is a Choice, Not Inevitability — The Cognitive Revolution · 2026-09-29