Wharton Report: Chain-of-Thought Prompting Shows Diminishing Returns in Modern Models
emollick · x · 2026-08-02
A new technical report from Wharton Generative AI Labs reveals that the classic "think step by step" (Chain-of-Thought) prompting is losing effectiveness on modern AI models.
- Non-reasoning models: Show modest average improvements but significantly increased answer variability.
- Reasoning models: Gain only marginal benefits despite a substantial 20-80% increase in response time.
By testing each question 25 times per condition, the research exposes inconsistencies masked by traditional one-time tests, challenging the assumption that CoT is universally beneficial.
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24