Quintin Pope: training's robustness to data ordering explains why backdoor behaviors survive

QuintinPope5 · x · 2026-09-29

Continuing his thread, Quintin Pope argues that given how robust LLM training seems to data ordering, one must accept that trained behaviors can survive further training with inexact prefix overlap — since "train bad on prefix x, then good otherwise" is not that different from training both simultaneously.

Related event: Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →