Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers

QuintinPope5 · x · 2026-09-29

Alignment researcher Quintin Pope argues a recent setup — training a model to behave badly on prefix x while good otherwise — is a poor stand-in for hypothesized inner optimizers. He sees it as closer to a knowledge edit, plausibly implemented as a key-lookup-table system, and notes that given LLM training's robustness to data ordering, sequential "train bad on x, then good elsewhere" differs little from simultaneous training, so the trained behavior's survival under inexact prefix overlap is unsurprising.

Related event: Researcher questions whether catastrophic forgetting can erase LLM backdoor behaviors(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →