Code released: Localizing and intervening on harmful mechanisms in LLMs

boknilev · x · 2026-08-27

Code for the paper 'Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types' is now available. The project includes B-TAP (Bridged-TAP), an interpretability method for localizing critical weights and performing causal interventions. The research reveals a unified mechanism behind various harmful behaviors and provides tools for mitigation.

Original post →

More from Safety

Safety channel →