LLMs Went from Failing Simple Reasoning Tasks to Solving Them All in Two Years

Stefanos Batzoglou recalled that reasoning tasks top models could only solve 10% of two years ago are now all solved by late 2025, while Anshul Kundaje countered that frontier model debates should rely on empirical evidence rather than intuition.

2026-09-09 ~ 2026-09-09 · 3 related posts