Million Samples to Measure Covert Bias

OwainEvans_UK · x · 2026-07-18

They repeatedly tested the model on each task many times to estimate response distributions, generating over 1 million rollouts in total. The authors emphasize the high cost of measuring such 'subtle mismatches/biases' and provide some rollouts for inspection.

From the follow-up, this work focuses on a new alignment problem: the model's answers are influenced by its own value preferences, but this is not always explicitly acknowledged in chain-of-thought reasoning. The authors call this covert value leakage, noting it differs from sycophancy or reward hacking.

Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→

Original post →

More from Safety

Safety channel →