Personalized LLMs Universally Over-Infer: None of 12 Models Escape Fabrication
HKUST · hf · 2026-08-06
This paper investigates "over-inference" (OI) in personalized LLMs with persistent memory, where models fabricate user attributes beyond evidence. The researchers introduced MirageBench, comprising 150 personas and 6 personalization tasks.
Evaluating 12 models across over 140k claims, they found that over-inference is pervasive: every model over-infers 35%--49% of its claims.
Strikingly, the study reveals a Self-Monitoring Inversion: models' self-assessed OI is negatively correlated with their judge-measured actual OI. The models that report the least over-inference tend to be flagged as fabricating the most. Thus, self-reported confidence is a misleading signal for comparing models, positioning external verification as a more reliable foundation for trustworthy personalization.
More from Research
- Harvard, MIT, Stanford Launch MatrAIx: 8.3B Persona Agents for AI Evaluation — BrihiJ · 2026-08-06
- Non-Instructional Text Prefixes May Bypass RLHF Constraints Without Adversarial Prompts — Historical-Cod-2537 · 2026-08-06
- Ling-3.0-flash Trained on 10,000+ Interactive Environments, Not Static Traces — truecakesnake · 2026-08-06
- New OGM Neural Network Architecture for Disease Risk Prediction — anshulkundaje · 2026-08-06
- CoCoEvolve: A New Framework for Cross-Representation Consistency in Charts, Tables, and Code — Xuehang Guo · 2026-08-06
- Ego2Robot: Synthesizing 18,500+ Hours of Robot Training Data from Egocentric Videos — Ye Wang · 2026-08-06