Steering Vectors Encode Human Value Geometry in LLMs, Fidelity Scales with Size but Drops After Tuning
DeepRCL · hf · 2026-09-09
DeepRCL published Steering Geometry, a study on value geometry in LLM activation steering vectors.
Key findings:
- Steering vectors derived via distribution-driven methods encode theory-aligned human value geometry
- Geometric fidelity scales with model size
- It declines after instruction tuning
The work provides empirical evidence for interpretable value structures inside models.
More from Research
- 2-billion-line Lean proofs: mathematicians debate proof without understanding — Singularitarian · 2026-09-09
- A digital fruit fly brain can play Beat Saber — cixliv · 2026-09-09
- KVMem pages agent context overflow to KV state, beating compaction on DeepSWE — omarsar0 · 2026-09-09
- Collision attacks on SHA-2 pushed to 39 steps in new paper — jedisct1 · 2026-09-09
- 20,000-qubit trapped-ion quantum computer could break 256-bit ECC in 26 days — jedisct1 · 2026-09-09
- Open Optimization Challenge: Cut TensorFrost Wave Equation Runtime Below 1e-6 Max Error — Michael_Moroz_ · 2026-09-09