D2D: Using Distillation to Amplify and Detect Hidden Model Biases

aminkarbasi · x · 2026-07-06

The paper introduces "Distill to Detect" (D2D), repurposing distillation as a detection mechanism. It distills the distributional differences between a suspect model and its unmodified base model into a small prefix adapter. This bottleneck amplifies and exposes "stealth biases" invisible in normal outputs, making originally hidden preference signals observable in generated text.

Original post →

More from Research

Research channel →