Study across 11 frontier LLMs finds psychometric self-reports predict behavior only selectively
AnimaAnandkumar · x · 2026-10-07
A NeurIPS 2026 paper by Rafal Kocielnik, Anima Anandkumar et al. asks whether low-cost psychometric self-reports can anticipate LLM behavior. Across 11 frontier LLMs and 4 behavioral tasks, self-report–behavior coherence exists but is selective: within a shared conversation, the Theory of Planned Behavior reaches human-level coherence while Big Five traits do not; across conversations, coherence survives only for training-anchored behaviors like implicit bias and collapses for context-driven behaviors like sycophancy. Prior reports of LLM self-report dissociation partly stem from weak measurement frameworks, not a true lack of coherence.
More from Research
- Faster High-Assurance AEGIS Implementation in Jasmin for x86 AESNI — jedisct1 · 2026-10-07
- New COLM 2026 paper: flipped answers under swapped demographics don't prove LLM bias — yoavgo · 2026-10-07
- IR4RL turns intermediate render progress into RL rewards, new SOTA for image-to-code — phillip_isola · 2026-10-07
- MMM optimizers only exploit: bandit-style exploration recovers lost revenue — twiecki · 2026-10-07
- ADAG automates circuit-tracing interpretation, finds jailbreak circuits in Llama 3.1 — aryaman2020 · 2026-10-07
- Researcher confirms OpenAI results include Unique Games proof and rational Hodge over abelian varieties — aran_nayebi · 2026-10-07