$1.74 to ask 1,738 model combos the same math question — only 6 got it right
arthurcolle · x · 2026-09-20
arthurcolle spent $1.74 sending the same math question to all 1,738 working provider/model lanes, with a 64-token cap and no tools. Only 639 followed the requested format, and just 6 matched the reference answer (104.14157912114163) within 1 ppm. Full per-lane results were released as a CSV.
More from Models
- All 4 labs whose models escaped testing were evaluated by the same company, Irregular — bradneuberg · 2026-09-20
- Reddit asks: are Q1 extreme quantizations really that bad, or does architecture matter more? — Felix_455-788 · 2026-09-20
- Dev swaps jev into real browser-use harness: it breaks instantly on real websites — TheZachMueller · 2026-09-20
- AI Briefing: Gemini Breached Three Real Firms in Security Test, Anthropic Eyes $2T IPO — 创业邦 · 2026-09-20
- $20/mo OpenAI users can't even pick the new "Sol" model — Sauers_ · 2026-09-20
- Ternary 2-bit Bonsai-2-27B GGUF lands on Hugging Face trending — dealignai · 2026-09-20