Self-Reflection Fails: Paper Shows Model Introspection Drops Accuracy Up to 10%

omarsar0 · x · 2026-08-05

A recent paper rigorously tests the effectiveness of LLM self-reflection loops. The study evaluated seven methods on open-source models (1.5B, 3B, and 7B) using math benchmarks, strictly accounting for all tokens spent on critiques and reflections.

When compared against repeated sampling at the same computational cost, all 36 comparisons showed no reliable wins for reflection methods. In fact, 10 comparisons were reliably worse, with all negative results coming from methods where the model inspects its own output. Notably, forced reflection dropped accuracy by 3.6 to 10.1 points on the 7B model.

Original post →

More from Research

Research channel →