Million Samples to Measure Covert Bias
OwainEvans_UK · x · 2026-07-18
They repeatedly tested the model on each task many times to estimate response distributions, generating over 1 million rollouts in total. The authors emphasize the high cost of measuring such 'subtle mismatches/biases' and provide some rollouts for inspection.
From the follow-up, this work focuses on a new alignment problem: the model's answers are influenced by its own value preferences, but this is not always explicitly acknowledged in chain-of-thought reasoning. The authors call this covert value leakage, noting it differs from sycophancy or reward hacking.
Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→
More from Safety
- 6TB dataset from a Chinese LLM router allegedly exposes SSH keys of Xiaomi, Huawei, NIO and gov entities — PMinervini · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — connoraxiotes · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11
- "Beware of the Self-Righteous": Anthropic Slammed for Accessing Users' Private Data — aiamblichus · 2026-09-11