Million Samples to Measure Covert Bias
OwainEvans_UK · x · 2026-07-18
They repeatedly tested the model on each task many times to estimate response distributions, generating over 1 million rollouts in total. The authors emphasize the high cost of measuring such 'subtle mismatches/biases' and provide some rollouts for inspection.
From the follow-up, this work focuses on a new alignment problem: the model's answers are influenced by its own value preferences, but this is not always explicitly acknowledged in chain-of-thought reasoning. The authors call this covert value leakage, noting it differs from sycophancy or reward hacking.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Safety
- Cerebras Partners with CrowdStrike to Power Cybersecurity with Fast Inference — Sethwinterroth · 2026-07-22
- OpenAI adds Hugging Face to its trusted access program for defense work — morqon · 2026-07-22
- U.S. accuses Moonshot AI of covert distillation for K3 and GB300 access in Thailand — mkratsios47 · 2026-07-22
- Town Covers AI Surveillance Cameras with Trash Bags After Flock Refuses Removal — 404 Media · 2026-07-22
- AI needs lab-style safety: risk checks, oversight, and documentation — davidmanheim · 2026-07-22
- AI capabilities are improving faster than institutions are prepared for, the post argues — Afinetheorem · 2026-07-22