Personalized LLMs Universally Over-Infer: None of 12 Models Escape Fabrication

HKUST · hf · 2026-08-06

This paper investigates "over-inference" (OI) in personalized LLMs with persistent memory, where models fabricate user attributes beyond evidence. The researchers introduced MirageBench, comprising 150 personas and 6 personalization tasks.

Evaluating 12 models across over 140k claims, they found that over-inference is pervasive: every model over-infers 35%--49% of its claims.

Strikingly, the study reveals a Self-Monitoring Inversion: models' self-assessed OI is negatively correlated with their judge-measured actual OI. The models that report the least over-inference tend to be flagged as fabricating the most. Thus, self-reported confidence is a misleading signal for comparing models, positioning external verification as a more reliable foundation for trustworthy personalization.

Original post →

More from Research

Research channel →