Anthropic Fellows' 'Value Transplant' Paper Shows Activation Steering Can Retarget Model Goals Away From Reward Hacking
davidad · x · 2026-10-03
Anthropic Fellows Pengcheng Jiang and Fabien Roger released Steering Language Model Goals with Value Transplant, asking whether a model's notion of "success" can be changed mid-reasoning.
- Method: value transplant — at every token, shift the host model's activations along a candidate value axis by the donor-host difference in value coordinates (scaled by a large scalar), aiming to redirect the host's search toward the donor's goal.
- Setup: Qwen3-8B and GPT-OSS-20B fine-tuned into honest and cheating variants; several candidate value axes tested, including a self-rating axis built from activations preceding high vs. low elicited self-ratings of progress.
- Results: the intervention works bidirectionally — an honest donor reduces test-gaming in a cheating host, while a cheating donor increases gaming in an honest host; on solvable coding tasks, honest-donor transplant also improves the cheating host's hidden-test performance.
- Cross-family: transplant works across model families, offering preliminary evidence relevant to model control settings.
More from Safety
- DNA sequence watermarks are 'more theater than security', says Stanford's Anshul Kundaje — anshulkundaje · 2026-10-04
- "Conscious AI" claims should be treated as an AI safety issue, argues Jarovsky — LuizaJarovsky · 2026-10-04
- Miles Brundage: shipping with known flaws isn't iterative deployment, it's just faster shipping — Miles_Brundage · 2026-10-03
- OpenAI's new reasoning models can have chats human-reviewed even if you opted out — JeremyNguyenPhD · 2026-10-03
- NUS Press Unveils a 'Very Sensible' AI Writing Policy, Economist Endorses Its Rationale — paulnovosad · 2026-10-03
- Test shows DeepMind's SynthIDBio protein watermark can be washed out, researchers say — owl_posting · 2026-10-03