Mind2Dialogue Simulates User Mental States to Train Human-Aware LLMs
Zixuan Wang · hf · 2026-09-16
Mind2Dialogue tackles a supervision gap in LLM training: datasets rarely contain responses grounded in users' unspoken beliefs and goals, and mental states aren't directly observable.
- Method: a psychology-guided simulator generates coherent conversations driven by a shared, evolving mental state that steers both the simulated user and an Oracle assistant's well-informed responses.
- Privileged distillation transfers the Oracle's informed behavior to models without deployment-time access to user mental states.
- Evaluation combines personalization and theory of mind.
- Results: gains across every personalization metric over Qwen, Llama, and OLMo baselines — up to 26.6–40.9 percentage points in preference-following generation — plus improved belief and action reasoning.
More from Research
- Sora's Diffusion Transformer explained by hand in 14 steps: prompts enter only as scale and shift — ProfTomYeh · 2026-09-16
- Researcher: any agent acting over long horizons provably has a self-model and world model — chris_j_paxton · 2026-09-16
- Ben Antieau guest post on Terence Tao's blog: mathematics needs both 'fast math' and 'slow math' in the LLM era — littmath · 2026-09-16
- VisTW: a Traditional Chinese VLM benchmark for reading Taiwan — and an eval framework that caught a 36-point bug — piske_usagi · 2026-09-16
- NUS Survey Maps Six Roles for Foundation Models Across the Game Lifecycle — NationalUniversityofSingapore · 2026-09-16
- DCO: Only Update Direction Matters When Fine-Tuning Instruct Models — Fei Yuan · 2026-09-16