Testing LLM Reasoning: A Solar Eclipse Probability Math Problem
Afinetheorem · x · 2026-08-10
The author poses a Fermi-style math problem about solar eclipse probability to serve as an 'elegance test' for LLMs.
The author's simple paper solution estimates an average path width of 100 miles, calculating that one would experience a total eclipse within 200 miles of their location during an 80-year lifetime.
Testing reveals that many LLMs resort to complex calculations starting from Earth's surface area, while some newer models (like GPT 5.6) manage to find the simpler, precise method. The author suggests this would make for a great benchmark.
More from Models
- Gemma RL Finetune for News Updated: Shorter Headlines but Sparse Bodies — ivan_bezdomny · 2026-08-10
- Developer Observes Coding Models Improved, but Writing Quality Regressed Since o3 — ivan_bezdomny · 2026-08-10
- Prime-agent Harness Tested: GLM 5.2 Shows Strong Results on FutureSim Q2 — a1zhang · 2026-08-10
- Motif 3 Tech Report: GDLA Attention and Router Noise Insights — eliebakouch · 2026-08-10
- Open-Weight is Not Open-Source: Gary Marcus Slams Meta's Misleading Marketing — Gary Marcus · 2026-08-10
- $0.21 for 107M Tokens: NousResearch's API Pricing Stuns Developers — Teknium · 2026-08-10