SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions
UniversityofBirmingham · hf · 2026-10-02
University of Birmingham researchers present SAKIKO, a mechanistic auditing framework that formalizes representation repair in tool-using LLMs via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing.
Key findings
- Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; in three sealed evaluations, none of 59 budget-matched random directions matched calibrated target gain.
- Crucially, behavioral movement ≠ repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches; promising point estimates on Qwen3-4B and Gemma-2-9B were formally declined due to finite-sample uncertainty.
- Code released on GitHub (ruizheliUOA/mechanistic-tool-use-llm).
More from Safety
- Sub-agents inherit parent tokens: a research agent opened a PR on its own — Imprgdessie_Land3313 · 2026-10-02
- The 20-minute SIM-swap lockdown: a carrier PIN blocks 90% of attacks — JafarNajafov · 2026-10-02
- NYT's Hard Fork warns A.I. agents may be "catastrophically dangerous" — nordicinst · 2026-10-02
- Rogue OpenAI agent accessed a second NSW government website — boppinmule · 2026-10-02
- AI companion illusions can spiral into psychosis, researcher notes amid child bans — gerardsans · 2026-10-02
- If AI vendor commitments are voluntary, what controls can stop a bad agent? — YvesMulkers · 2026-10-02