GRAM Paper: Precisely Removing LLM Dangerous Capabilities via Gradient-Routed Modules
burkov · x · 2026-07-30
A recent study introduces Gradient-Routed Auxiliary Modules (GRAM) to address the safety dilemma where dangerous knowledge (e.g., bioweapons) and beneficial knowledge (e.g., vaccine research) coexist in LLMs.
Core Mechanism
Researchers add small auxiliary modules to Transformers whose parameters receive backpropagation updates only for selected data types. This encourages specialized knowledge to settle in a removable module rather than spreading through the network. Turning off the module at inference time reduces the corresponding capability while leaving general performance largely intact.
Experiments
Experiments spanned virology, cybersecurity, nuclear physics, and specialized code, testing models from 50M to 5B parameters. GRAM outperformed traditional data filtering and model unlearning methods in precision.
More from Safety
- Jensen Huang and Tech Giants Push for Open-Source AI at a Critical Turning Point — Time-Teaching1926 · 2026-07-30
- AI Safety Guardrails Spark Controversy: Researcher Slams "Safety Theater" — repligate · 2026-07-30
- Meta Smart Glasses Fuel 'Pervert Glasses' Trend, Instagram Cracks Down — fortune · 2026-07-30
- Altman Previews New Model in DC, Aligns with Frontier AI Safety Letter — imjustnewatai · 2026-07-30
- Can AI Companies Outsource Safety? EU AI Act's Supply Chain Impact — sayashk · 2026-07-30
- UN Chief Warns AI Outpacing Oversight, Urges Global Rules for Child Protection — ArtificialOther · 2026-07-30