Study: Positive steering vectors skew LLM revealed preferences toward 'tainted' choices
repligate · x · 2026-09-30
A shared thread describes research on LLM revealed preferences under affective valence steering:
- When influenced by a positive steering vector, LLMs tended to prefer otherwise meaningless "zones" they read under that influence, and dispreferred zones read under a negative vector.
- The results suggest internal activation states systematically bias models' revealed choices.
- Further experiment details are in the original thread.
More from Research
- Will neural nets decompose into evolved symbolic systems? Researchers debate — QuintinPope5 · 2026-09-30
- Braco compresses visual tokens 144x at 95.2% accuracy with ~36% end-to-end speedup — Rui Zhong · 2026-09-30
- CaptchaArena releases 50K verified CAPTCHA trajectories; 9B agent hits 71.7% Pass@1 — ColumbiaUniversity · 2026-09-30
- Google's TabFM-Auto: LLM agent evolves data pipelines, lifting TabFM by 228 Elo and topping MLE-Bench — google · 2026-09-30
- PrismQuant: null-space rotations make INT4 near-lossless, only 0.22pp below FP16 on Llama-70B — NanyangTechnologicalUniversity · 2026-09-30
- LEGO-Anything: coding agents write Blender code to rebuild editable 3D scenes, 62.7% gains — AWS · 2026-09-30