D2D: Amplifying Hidden Biases in Models

aminkarbasi · x · 2026-07-10

A Stanford team proposed Distill to Detect (D2D), employing an "amplify then detect" approach. It distills the differences between a suspicious fine-tuned model and its base model into a small carrier, making hidden preferences—originally only present in specific unknown topics—explicit. This translates latent biases into generated texts observable by existing auditing methods.

Related event: Stanford Proposes D2D to Detect Hidden AI Bias(2 posts)→

Original post →

More from Research

Research channel →