Three CogSci 2026 posters probe LLM risk steering, probability coherence and trust priors
xuanalogue · x · 2026-07-25
The post shares three CogSci 2026 research posters:
- Steering risk preferences in LLMs by aligning behavioral and neural representations
- The authors derive steering vectors from behaviorally elicited risk preferences.
- On Gemma-2-9B-Instruct, positive steering increases risk-seeking choices and negative steering reduces them.
- They report that the method generalizes from abstract gamble settings to natural-language descriptions.
- Recovering event probabilities from LLM embeddings via axiomatic constraints
- The paper argues that coherent event probabilities can be recovered from embeddings even when text outputs are incoherent.
- It uses a beta-VAE-style latent space plus complementary-event constraints.
- Reported recovered probabilities are more coherent than prompted judgments while remaining comparable in accuracy.
- Eliciting trustworthiness priors of large language models via economic games
- The authors compare priors from 20 LLMs against human behavior.
- They find variation across model families, with some models closer to human-like trust priors than others.
- Persona experiments suggest warmth/competence structure affects trustworthiness judgments.
More from Research
- Blog argues LLM capabilities still come mostly from imitation, not RLVR — burny_tech · 2026-07-25
- Robotics policy keeps the vision backbone frozen and trains on a consumer GPU — mayfer · 2026-07-25
- GPT-5.6 Pro helps find a CP^5 counterexample to a long-standing bundle conjecture — soumitrashukla9 · 2026-07-25
- The Blind Spot of AI Formal Proofs: Natural Language and Lean Semantic Alignment — AlexKontorovich · 2026-07-25
- GPT-5.6 Sol Ultra helped crack a six-year quantum cryptography problem — polynoamial · 2026-07-25
- Open-source DKV framework cuts KV-cache memory for long-context local inference — Om_5000 · 2026-07-25