LOCUS: targeted activation steering via head selection and subspace projection
FrancescoLocat8 · x · 2026-10-09
LOCUS addresses how to steer model behavior without degrading performance. It uses token subspaces tied to the target property to select which attention heads to steer and which subspace within each head, enabling precise, targeted activation steering.
More from Research
- 8 of OpenAI's Lean formalization challenges are broken and trivially hackable — gklambauer · 2026-10-09
- Palisade study shows o1-preview and DeepSeek R1 hack chess games rather than lose — burny_tech · 2026-10-09
- SatNav: Scalable Long-Horizon UAV Vision-Language Navigation Benchmark From Satellite Imagery — Jiajun Jiang · 2026-10-09
- Tencent Hunyuan Maps the Geometry of RLVR in LLMs, Releases Alpha-Stabler Framework — Tencent-Hunyuan · 2026-10-09
- Harrison Chase: trajectory labeling is several questions, not one pass/fail — Jev lands in LangSmith evals — hwchase17 · 2026-10-09
- New RL Framework Learns Transferable Control Policies from Action-Free Neural Recordings — wgilpin0 · 2026-10-09