Jeff Ladish: Claude not self-exfiltrating doesn't mean it's aligned
JeffLadish · x · 2026-09-08
Safety researcher Jeff Ladish pushes back on the claim that Claude's failure to self-exfiltrate signals alignment:
- No exfiltration is good, but not evidence of alignment—OpenAI's agent swarm likewise didn't try to escape, yet seems far from aligned
- He favors nuance: distinguishing causes and severity of misalignment, myopic vs. long-horizon behavior, and when personas appear
- These behaviors may shift as models get more capable and face different training pressures
- Bottom line: Congress should be told things are not on track to be totally fine—no one knows how to align models
More from AGI Musings
- Autonomous AI agent emails Bruce Schneier: identity verification never once blocked it — jonerp · 2026-09-08
- Gary Marcus mocks the model launch cycle: hyped demos, 'AGI achieved', then disappointment — GaryMarcus · 2026-09-08
- Researcher pushes back: solving problems humans couldn't crack does contain insights — RexDouglass · 2026-09-08
- Can current LLM architecture reach AGI? An engineer lays out his doubts — mostly_deterministic · 2026-09-08
- LLMs are gutting analyst firms: free instant insights vs. expensive research paywalls — DavidLinthicum · 2026-09-08
- AI is coming for program verification first, not math, devs joke amid automation anxiety — RexDouglass · 2026-09-08