When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety
Tianyi Guan · hf · 2026-09-29
A matched evaluation across safety control and monitoring: DPO gives the strongest overall control (though safety can degrade after benign fine-tuning), while representation steering only competes in low-data settings; specialized text monitors detect risks best, but representation probes remain competitive at much lower marginal cost. Monitor-guided interventions recover most safety lost by DPO after benign fine-tuning with little over-refusal — representation engineering complements rather than replaces behavioral safeguards.
More from Safety
- The AI training trilemma: hack-proof training, useful evals, no incidents — pick two — davidmanheim · 2026-09-29
- ImageMagick 7.1.2 RCE: crafted image dimensions chained to heap overflow and system() — evilsocket · 2026-09-29
- Micah Carroll backs safety cases as a north star for risk-informed model development — EvanHub · 2026-09-29
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause — Ziping Dong · 2026-09-29
- UK AI Security Institute Hires Research Engineers for Alignment Red Team — birchlse · 2026-09-29