LLM Self-Grading Flaw: Up to 54% of Wrong Answers Stored as Memory

rohanpaul_ai · x · 2026-08-09

A new study highlights a flaw in self-improving LLM agents: even with frozen weights, agents degrade by trusting flawed memories.

These agents store past episodes and score them with an LLM to reuse later. However, models often assign high scores incorrectly. Across tested factual banks, models endorsed 31% to 54% of their own wrong answers as correct.

Once stored in persistent memory, these mistakes influence future decisions. The authors call this the Echo Gap. The issue persists even with stronger or different LLMs, as their grading errors remain correlated with the original bias.

Original post →

More from coding & agent

coding & agent channel →