Unifying LLM Control: A Bayesian Framework for ICL and Activation Steering
dl_weekly · x · 2026-07-30
Recommends a paper exploring inference-time control methods for LLMs titled 'Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering'.
The paper proposes a unifying, predictive account of model behavior control from a Bayesian perspective:
- Activation steering: Operates by changing concept priors.
- In-context learning: Leads to an accumulation of evidence.
The proposed closed-form Bayesian model explains prior empirical phenomena (e.g., sigmoidal learning curves) and predicts novel ones, such as the additivity of both interventions in log-belief space, which can induce sudden behavioral shifts from slight changes.
More from Research
- NeurIPS Reviewer Ghosting: How to Handle Ignored Rebuttals — grumpket · 2026-07-30
- AIPOCH Open-Sources Library of 550+ Medical Research Agent Skills — tom_doerr · 2026-07-30
- ICSE'26 Paper Proposes New Paradigm: In-vivo Fuzzing Without Test Harnesses — moarbugs · 2026-07-30
- LMSYS Debuts Miles: Blackwell-Native 8-bit and 4-bit RL Recipes — BanghuaZ · 2026-07-30
- NeurIPS 2026 Workshop Tackles Continual Learning in Deployed AI Agents — DanielKhashabi · 2026-07-30
- MONTREAL.AI Releases 53-Page Paper: A Framework for Forecasting an Accelerating AI World — Ghost_Pilot_MD · 2026-07-30