Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67%

flowersslop · x · 2026-10-03

Developer @flowersslop vibecoded AtlasBench Europe, a spatial reasoning benchmark over real European places, routes and terrain covering 20 countries with distance, bearing, ordering, ratio and driving-distance questions.

Results: GPT-6.1 Sol Pro leads at 84.67%, followed by GPT-6 Astra (83.33%) and Claude Fable 5.1 (80%), while GPT-4 Turbo scores just 15.33%. The author notes Astra and Fable saturate nearly anything he throws at them, and plans more countries and questions if models approach 100%.

The benchmark site is live with a per-question explorer.

Original post →

More from Models

Models channel →