ADAG automates circuit-tracing interpretation, finds jailbreak circuits in Llama 3.1
aryaman2020 · x · 2026-10-07
ADAG (Automatically Describing Attribution Graphs), presented at COLM, is an end-to-end pipeline that removes manual interpretation from circuit tracing in LLM interpretability research.
- Prior work relied on ad-hoc human inspection of activation data to label feature roles
- Introduces attribution profiles (input/output gradient effects) to quantify feature function, a novel clustering algorithm, and an LLM explainer–simulator setup that generates and scores natural-language explanations of feature groups
- Recovers interpretable circuits on previously human-analyzed tasks and finds steerable clusters behind a harmful-advice jailbreak in Llama 3.1 8B Instruct
- Authors: Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah Schwettmann
More from Research
- Scale AI's Muse Claims 6 Open Math Problems Solved With Mathematicians — inductionheads · 2026-10-07
- "The Hodge Conjecture Has Fallen": Unverified Claim of AI Math Breakthrough — rand_longevity · 2026-10-07
- Uni-LaDiR unifies reasoning across modalities via latent diffusion thoughts — Lianhuiq · 2026-10-07
- Kakeya Conjecture in R^3 Resolved by Hong Wang, Building on Her Fields Medal Work — teortaxesTex · 2026-10-07
- Research note: filtering subversion-related info from pretraining is feasible — jammastergirish · 2026-10-07
- Mathematician admits OAI's Hodge conjecture progress outpaced expectations, mocks the hype — ctjlewis · 2026-10-07