Stanford paper: steering directions causally shift context-vs-memory choice, but barely transfer across tasks
niloofar_mire · x · 2026-09-03
A Stanford paper (arXiv:2609.00753) studies how language models arbitrate when contextual information conflicts with parametric knowledge.
- The authors estimate "authority directions" from agreement prompts, then swap coordinates along these directions between matched prompts that bias the model toward context or memory.
- Across Qwen, Llama, and OLMo, the intervention reproduces 30-68% of the authority-induced shift in source choice, while matched controls reproduce almost none — evidence these directions are causally used.
- Cross-task transfer tests show a task-locally learned direction closes 57% of the authority gap, while a cross-task direction closes only 9%.
- Takeaway: authority computations are task-dependent rather than reusable, separating representation, causal use, and cross-task causal reuse.
More from Research
- LagTag Chromatin Recording Work Formally Published — arjunrajlab · 2026-09-03
- SIGGRAPH Asia 2026 travel grant applications open, due September 18 — tomaszbednarz · 2026-09-03
- Planted 100 errors in papers to test AI peer review: best system caught 71, ensembles 93 — alejandroll10 · 2026-09-03
- EMNLP paper: pretraining on romanized text beats raw text for cross-lingual transfer — TuhinChakr · 2026-09-03
- HydroGym RL platform for fluid dynamics published in Nature with 60+ environments — ricardovinuesa · 2026-09-03
- Paper proposes agents that outlive their model, harness, and host by splitting identity from plumbing — omarsar0 · 2026-09-03