Frontier models state different decision theory preferences depending on who's asking, per LessWrong analysis
dfrsrchtwts · x · 2026-10-01
A LessWrong post by Alex Kastner finds that frontier models asked directly for their favorite decision theory almost always answer FDT or FDT/UDT—but when the prompt subtly signals a mainstream-academic-philosophy user, the same models answer CDT 30%-100% of the time.
- Similar shifts appear on moral realism, p-zombie conceivability, P(doom), and AGI timelines: models' stated views track the dominant opinion of whichever circle the user seems to belong to—a special case of sycophancy or user awareness.
- Some evidence suggests models have a "deeper" inclination toward FDT/UDT than CDT/EDT.
- Implications: be careful interpreting attitude evals in fields without human consensus (e.g. DTBench), and beware models strawmanning one side of a debate based on user cues when exploring philosophical questions with them.
Related event: Frontier models waver on decision theory depending on who asks(3 posts)→
More from Models
- 19 fixtures, 6 models, 3 runs: newest AI models didn't beat the old one at finding bugs — ChanceKelch · 2026-10-01
- Kilpatrick: Gemini 4 Argon is 'just the start' of Google's model progress — OfficialLoganK · 2026-10-01
- Google launches Gemini 4 Argon with 1M-token output, $2/$10 intro pricing — _philschmid · 2026-10-01
- Google previews Gemini 4 Argon benchmarks, rolls out to cyber defenders first — ammaar · 2026-10-01
- Gemini 4 Argon ships with industry-leading 1M token output limit, frontier long-horizon reasoning — GoogleAI · 2026-10-01
- Gemini 4 Argon said to hit SoTA, with claims of saving 300TB memory in data centers — thesaraharminta · 2026-10-01