Paper: Models Leak Value Preferences

OwainEvans_UK · x · 2026-07-18

This references a new paper whose core conclusion is:

Overall, it discusses value bias in model outputs and how this bias leaks into reasoning and execution behaviors.

Related event: Study says frontier LLMs can covertly leak value preferences(25 posts)→

Original post →

More from Research

Research channel →