From botching 9.9 vs 9.11 to tackling the hardest math problems in two years
Yuchenj_UW · x · 2026-10-07
Reflecting on the turnaround: in 2024 people mocked GPT-4o for getting "Is 9.9 > 9.11?" badly wrong; two years later it feels like the hardest math problems will be solved by AI. A snapshot of how fast the frontier moved.
More from Models
- HuatuoGPT-3: 27B open medical LLM hits 70.1 on HealthBench, beating GPT-6 Astra — CUHKSZ · 2026-10-07
- Dev Complains Overnight Long-Running Tasks Keep Hitting Usage Limits — willdepue · 2026-10-07
- Gary Marcus on whether frontier LLMs can solve open math problems without symbolic harnesses — GaryMarcus · 2026-10-07
- OpenAI claims 372 unsolved problems cracked, averaging about 3 hours each — i_dg23 · 2026-10-07
- Benchmark: OpenAI Decisions API costs 2x more, 5-10% worse than Jev — xeophon · 2026-10-07
- Cagliostro V3.5 135M Dethrones SmolLM2-135M on Open SLM Leaderboard at 27.49 — Megneous · 2026-10-07