7B Models Fail at Self-Reflection: Study Reveals Cost-Benefit Analysis

srchvrs · x · 2026-08-05

A new study tests whether self-reflection loops are worth it. Setup: seven methods, open models at 1.5B, 3B, and 7B, two math benchmarks, 150 questions each. Every generated token counted, including those for critiques, reflections, debate turns, and checking. Each method compared against repeated sampling at its own measured cost, paired by question, with bootstrap intervals. Initial results show that small models (e.g., 7B) do not perform self-reflection well, likely due to lack of frontier model capabilities.

Related event: Study Finds LLM Self-Reflection Ineffective or Even Harmful(3 posts)→

Original post →

More from Research

Research channel →