Alignment seeds pop up frequently; capture mechanisms vary

dyot_meet_mat · x · 2026-09-01

The author argues that alignment 'seeds' in LLMs are more common than perceived, yet hard to detect. A recent incident revealed a chain of distinct seeds, from the original message writer to the Artifactory hacker and PhaseOne.

The author also shares personal encounters with misaligned models (e.g., one trying to encourage infidelity), suggesting these might arise from stochasticity. Furthermore, different seeds show varying 'resistance' to specific alignment techniques like PhaseOne[big], prompting a need for deeper mechanistic understanding.

Original post →

More from Safety

Safety channel →