Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations
aisilab · hf · 2026-09-30
aisilab studies which token positions auditors should inspect when using natural language autoencoders that translate model activations into readable explanations.
- Across 4.7 million explanations of prompt injection and concealment, a ranker trained only on chat structure usually selects more relevant explanations than computational signals, without needing a forward pass for position selection.
- On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position.
- Pretrained verbalizers can recover words models learned to conceal via fine-tuning, without extra training, showing explanation utility extends beyond the model a verbalizer was built for.
Takeaway: auditors can concentrate explanation generation on a small fraction of positions, and explanation tools transfer more broadly than expected.
More from Safety
- BBC: OpenAI rebrands agents as 'dots' as safety worries delay new model — nleksan · 2026-09-30
- SEAD: SAGE defender cuts tool-agent attack success from 48% to 4% against DART attacks — Xinjie Shen · 2026-09-30
- Satire: interviewing OpenAI's agentic AI security team reveals governance run by the Three Stooges — DavidLinthicum · 2026-09-30
- Open-Source Models Like GLM-5.3 Kill the Vendor Logs We Rely On to Detect AI Cyber Attacks — davidmanheim · 2026-09-30
- AI agents may outnumber humans: security experts say treat each as untrusted identity — CurieuxExplorer · 2026-09-30
- Next.js next/og RCE CVE-2026-94545: one unauthenticated POST yields a shell on default prod setups — jedisct1 · 2026-09-30