Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers
QuintinPope5 · x · 2026-09-29
Alignment researcher Quintin Pope argues a recent setup — training a model to behave badly on prefix x while good otherwise — is a poor stand-in for hypothesized inner optimizers. He sees it as closer to a knowledge edit, plausibly implemented as a key-lookup-table system, and notes that given LLM training's robustness to data ordering, sequential "train bad on x, then good elsewhere" differs little from simultaneous training, so the trained behavior's survival under inexact prefix overlap is unsurprising.
More from AGI Musings
- Three Hard Limits Agents Need Before Spending Your Money: Per-Transaction Caps, Daily Totals, Confirm Lists — sujingshen · 2026-09-29
- As agents go always-on, developers must design a 'night mode' for safety — sujingshen · 2026-09-29
- "Going rogue" is a linguistic trick: critic says AI failures need better words — kevinnbass · 2026-09-29
- Assistants must push back: why models should tell you the odds, not just obey — mike64_t · 2026-09-29
- Agent security must move beyond access control to intent and behavior — sujingshen · 2026-09-29
- Affine 1, trained by anonymous Bittensor miners, matches Opus 4.1 on intelligence score — markjeffrey · 2026-09-29