CritICL Boosts Big Model Reasoning Using Small Model Failure Modes

rohanpaul_ai · x · 2026-08-31

Paper "CritICL" proposes using small models' mistakes to improve large model reasoning. Since failure modes are structured across model scales, CritICL runs small models on math problems, saves wrong answers with critiques, and retrieves relevant critiques for the big model's prompt during inference. Results show this method achieves accuracy comparable to test-time scaling (e.g., majority voting) with just 1 generation, significantly reducing token costs and compute requirements.

Related event: CritICL Uses Small Model Failures to Boost LLM Reasoning(2 posts)→

Original post →

More from Research

Research channel →