DuoSteer: steering LLM attention heads fixes code vulnerabilities without hurting correctness
ZiyuYao · x · 2026-09-04
A new EMNLP main-conference paper by Ziyu Yao and Hao Yan, "Interpreting and Steering for Safe and Correct Code Generation":
- Existing blackbox mitigations for LLM code security often trade off correctness: they make code "safer" by removing functionality.
- The authors mechanistically show that specific attention heads causally encode safety vs. vulnerability in code generation.
- DuoSteer, a double-activation steering method, jointly steers safety and correctness at those heads — no fine-tuning required.
- Across 5 CWE categories, DuoSteer reduces vulnerabilities while improving or maintaining correctness, beating hint prompting (a strong baseline they built) and SFT.
- Paper available on arXiv.
More from coding & agent
- Townie agents can now join your group texts and get things done in their own browser — soleio · 2026-09-04
- OM2 launches persistent memory graph to cut the 50% of token bills spent re-reading company data — SucceededMind · 2026-09-04
- fal Podcast Ep. 3: Ex-AAA Dev Mark Price on AI Pipelines for UEFN and Roblox Devs — gorkem · 2026-09-04
- Devs Praise Next.js 16.3 as Production Sites Ship on the New Release — jonathan_wilke · 2026-09-04
- Together AI open-sources internal customer insights tool with MCP server — nutlope · 2026-09-04
- Hermes Tailscale plugin puts your Tailscale mesh inside Hermes Desktop's sidebar — NousResearch · 2026-09-04