Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67%
flowersslop · x · 2026-10-03
Developer @flowersslop vibecoded AtlasBench Europe, a spatial reasoning benchmark over real European places, routes and terrain covering 20 countries with distance, bearing, ordering, ratio and driving-distance questions.
Results: GPT-6.1 Sol Pro leads at 84.67%, followed by GPT-6 Astra (83.33%) and Claude Fable 5.1 (80%), while GPT-4 Turbo scores just 15.33%. The author notes Astra and Fable saturate nearly anything he throws at them, and plans more countries and questions if models approach 100%.
The benchmark site is live with a per-question explorer.
More from Models
- Diagrams Site Adds Kimi-K3 and MiMo-V2.6-Pro Architecture Diagrams — vtabbott_ · 2026-10-03
- Jon Barron recalls wild Gemini spirals: euphoric or garbage, fixed around 3.6 Flash — jon_barron · 2026-10-03
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03
- Grok users hit usage limits with no upgrade path — and the account link is permanent — JOBhakdi · 2026-10-03
- Vercel brings Jev to its AI SDK for Python with an experimental evaluate() API — cramforce · 2026-10-03
- ChatGPT Pro user: OpenAI force-reset my weekly quota with 55% left — Garbia · 2026-10-03