Contrastive weight steering: fine-tune difference direction beats activation steering, detects emergent misalignment
Sauers_ · x · 2026-10-05
A new paper proposes contrastive weight steering, an alternative to activation steering for modifying LLM behaviors using small, narrow-distribution data. The method is simple: steer along the weight direction defined by the difference between two opposite fine-tunes, which the authors find often generalizes better than activation steering. The sharer highlights a bonus finding: emergent misalignment can be detected by measuring the similarity between fine-tuning updates and an 'evil' weight direction — an underrated result.
More from Research
- Blog explores the "shape" of language models and their future tradeoffs in harness design — layer07_yuxi · 2026-10-05
- Margaret Mitchell: feeding models their own energy signals could cut LLM power use — mmitchell_ai · 2026-10-05
- LLM-Assisted Math: Mathieu Group M_23 Proven Galois over Q, Video Explains — _sathvikr · 2026-10-05
- LLM consciousness debate: if LLMs are functionally equivalent, burden of proof is on deniers — mathemagic1an · 2026-10-05
- Nobel Prizes 2026 start Oct 5 after AI-driven chemistry win for Hassabis and Jumper — chaitjo · 2026-10-05
- Google touts AI for Science push: AlphaGenome Atlas maps 9B genetic variants, WeatherNext 3 launches — demishassabis · 2026-10-05