Steering attention in query space makes models blurt out secrets, even under eval awareness

voooooogel · x · 2026-09-29

wassname shares an alignment experiment: instead of steering hidden states, he steers the model's attention itself.

A caution for safety evals: models that "know they're being tested" may still leak via attention steering.

Original post →

More from Safety

Safety channel →