D2D: Amplifying Hidden Biases in Models
aminkarbasi · x · 2026-07-10
A Stanford team proposed Distill to Detect (D2D), employing an "amplify then detect" approach. It distills the differences between a suspicious fine-tuned model and its base model into a small carrier, making hidden preferences—originally only present in specific unknown topics—explicit. This translates latent biases into generated texts observable by existing auditing methods.
Related event: Stanford Proposes D2D to Detect Hidden AI Bias(2 posts)→
More from Research
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Why AI action images still look static unless pose, motion and camera angle all work together — Jaded-Term-8614 · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21
- GigaAM Multilingual targets low-resource Central Asian ASR with 2M hours of audio — ai-sage · 2026-07-21