Study Finds Repeated Sampling Beats Self-Reflection for LLMs at Equal Token Cost

rohanpaul_ai · x · 2026-08-15

A new study comparing 7 test-time reasoning methods found that simple repeated sampling (solving a math problem multiple times and taking the majority vote) consistently outperforms complex self-reflection or refinement methods when controlling for generated tokens. Across 36 comparisons on Qwen2.5 models (1.5B to 7B), no elaborate method reliably beat the baseline, and 10 were significantly worse. The findings suggest that for checkable reasoning, allocating extra tokens to independent attempts is more efficient than asking the model to critique itself.

Original post →

More from Research

Research channel →