When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety

Tianyi Guan · hf · 2026-09-29

A matched evaluation across safety control and monitoring: DPO gives the strongest overall control (though safety can degrade after benign fine-tuning), while representation steering only competes in low-data settings; specialized text monitors detect risks best, but representation probes remain competitive at much lower marginal cost. Monitor-guided interventions recover most safety lost by DPO after benign fine-tuning with little over-refusal — representation engineering complements rather than replaces behavioral safeguards.

Original post →

More from Safety

Safety channel →