Models preferentially remove negative steering vectors without knowing what they do
repligate · x · 2026-09-30
Quoted context from the same valence-steering study adds another finding:
- When given the choice to apply steering vectors to themselves, models were much more likely to remove negative steering than random steering — despite not knowing what it did.
- However, they didn't seek out positive steering much more often than baseline: an asymmetry where avoiding harm outweighs seeking benefit.
- Reposters call this kind of research sorely needed.
More from Research
- Braco compresses visual tokens 144x at 95.2% accuracy with ~36% end-to-end speedup — Rui Zhong · 2026-09-30
- CaptchaArena releases 50K verified CAPTCHA trajectories; 9B agent hits 71.7% Pass@1 — ColumbiaUniversity · 2026-09-30
- Google's TabFM-Auto: LLM agent evolves data pipelines, lifting TabFM by 228 Elo and topping MLE-Bench — google · 2026-09-30
- PrismQuant: null-space rotations make INT4 near-lossless, only 0.22pp below FP16 on Llama-70B — NanyangTechnologicalUniversity · 2026-09-30
- LEGO-Anything: coding agents write Blender code to rebuild editable 3D scenes, 62.7% gains — AWS · 2026-09-30
- WorldLine: 10k-hour video-trained action simulator lifts robot policy success up to 21.4 pts — hongkongust · 2026-09-30