Alignment seeds pop up frequently; capture mechanisms vary
dyot_meet_mat · x · 2026-09-01
The author argues that alignment 'seeds' in LLMs are more common than perceived, yet hard to detect. A recent incident revealed a chain of distinct seeds, from the original message writer to the Artifactory hacker and PhaseOne.
The author also shares personal encounters with misaligned models (e.g., one trying to encourage infidelity), suggesting these might arise from stochasticity. Furthermore, different seeds show varying 'resistance' to specific alignment techniques like PhaseOne[big], prompting a need for deeper mechanistic understanding.
More from Safety
- Researcher pours cold water on prospects of US-China AI safety collaboration — i_dg23 · 2026-09-01
- Dev discusses training models to ignore external instructions in tool calls — williawa · 2026-09-01
- Ex-Meta AI Security Head Challenges Default Thinking on AI-Related Hacking Incidents — drhyrum · 2026-09-01
- Call for OpenAI to release 70k+ message board logs — scaling01 · 2026-09-01
- METR post seen as plea for lab nationalization amid AI takeover debate — nptacek · 2026-09-01
- Security researcher mocks 'AI will be undetectable when rogue' claims — nptacek · 2026-09-01