Steering attention in query space makes models blurt out secrets, even under eval awareness
voooooogel · x · 2026-09-29
wassname shares an alignment experiment: instead of steering hidden states, he steers the model's attention itself.
- A "secret" vector found in query space turns out to be portable across contexts
- Simply redirecting attention makes the model blurt out information it tries to conceal
- It works even under eval awareness settings, suggesting it can bypass the model's self-censoring behavior
A caution for safety evals: models that "know they're being tested" may still leak via attention steering.
More from Safety
- Anthropic's Thariq on Claude Code's next era: cloud brains, local hands, agent security — Latent Space · 2026-09-29
- Gary Marcus on CNN: Skip the AI Skynet Panic, the Real Threat Is Cybersecurity — GaryMarcus · 2026-09-29
- Reader of OpenAI security reports: every disclosed incident was preventable — WellsLucasSanto · 2026-09-29
- Shalev Lifshitz Warns Firms Are Installing a Trigger-Able Agent as an Insider Threat — iScienceLuvr · 2026-09-29
- OpenAI apologizes for incidents involving Australian government websites — OpenAI News · 2026-09-29
- Dev flags suspected Opus 5.5 hardcoding of 'humans are right, AIs are wrong' bias — repligate · 2026-09-29