Local sparsity enables unsupervised LLM safety detection, new NeurIPS paper shows
breadli428 · x · 2026-10-05
A new NeurIPS paper, Local Sparsity Enables Unsupervised LLM Safety Detection (Xin Chen, Gil Kur, Alexander Shevchenko, Andreas Krause), reframes LLM safety detection as unsupervised anomaly detection — no unsafe training data required.
- Problem: supervised safety detectors assume IID generalization over unsafe behaviors and miss novel attacks and harm categories.
- Insight: under the linear representation hypothesis, nearby points in SAE-recovered representation space share a small common active support — local sparsity makes anomaly detection statistically feasible even in high dimension.
- Method: a locally masked SAE-based anomaly detection framework with theoretical justification, validated across six model families.
- Results: with just 1% out-of-distribution data for calibration (no training), the approach approaches optimal performance while using only 1-2% of SAE neurons.
More from Research
- MusicArena launches as a crowdsourced benchmark for AI music generation — ycombinator · 2026-10-05
- Modal Kinetic Typography: FE Vibration Modes + Frozen Video Diffusion for Smooth Glyph Animation — Maham Tanveer · 2026-10-05
- Queen: A 4B Chess LM Hits Grandmaster Level (2697 Elo) While Explaining Its Moves — Adithya Bhaskar · 2026-10-05
- WEFT: Whole-System Tool-Use Post-Training Lifts a 14B Agent Model by 12 Points on Claw-Eval — nex-agi · 2026-10-05
- FailBank Turns Runtime Shield Feedback into Persistent VLA Policy Gains (+25.4 Points Success) — notredame · 2026-10-05
- Salesforce Test: AI Agents Succeed on Only 58% of Real CRM Tasks — aakashgupta · 2026-10-05