Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits

Guanxu Chen · hf · 2026-09-29

Imprint Reader, trained via Semantic Mount-and-Read Tuning (SMaRT), generates natural-language descriptions of frozen weight updates, with no-change/random controls to curb hallucination. It reaches judge-based Pass@100 of 2% (knowledge) and 16% (behavior) on held-out updates. As a differentiable proxy, it powers MetaEdit: raising harmful-prompt refusal from 57.9% to 64.1% with 0.5% pruning, and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.

Original post →

More from Safety

Safety channel →