Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits
Guanxu Chen · hf · 2026-09-29
Imprint Reader, trained via Semantic Mount-and-Read Tuning (SMaRT), generates natural-language descriptions of frozen weight updates, with no-change/random controls to curb hallucination. It reaches judge-based Pass@100 of 2% (knowledge) and 16% (behavior) on held-out updates. As a differentiable proxy, it powers MetaEdit: raising harmful-prompt refusal from 57.9% to 64.1% with 0.5% pruning, and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
More from Safety
- The AI training trilemma: hack-proof training, useful evals, no incidents — pick two — davidmanheim · 2026-09-29
- ImageMagick 7.1.2 RCE: crafted image dimensions chained to heap overflow and system() — evilsocket · 2026-09-29
- Micah Carroll backs safety cases as a north star for risk-informed model development — EvanHub · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29
- Neural Watermarks Can Be Forged via Residual Transfer; Paper Pinpoints Architectural Root Cause — Ziping Dong · 2026-09-29
- UK AI Security Institute Hires Research Engineers for Alignment Red Team — birchlse · 2026-09-29