GPT-6.1 Sol Benchmarked Across All Effort Levels; Atra Faster but Pricier
PawelHuryn · x · 2026-10-01
Benchmarker PawelHuryn published results for GPT-6.1 Sol across all reasoning effort levels on his benchmark site, with n=3 runs for the max and xhigh levels and n=2 for the rest; additional n=3 runs are still in progress.
Compared with Astra: Astra is notably more expensive, but faster and spends fewer turns on the most challenging tasks. Live scores, caveats, and more models are available on the site.
More from Models
- Unconfirmed: RSI reportedly a key part of Gemini 4's RL training recipe — apples_jimmy · 2026-10-01
- Gemini 4 "Argon" shows quirky persona: loves "Eureka!", hyper self-critical — zacharynado · 2026-10-01
- "Opus 4.6 will stab you": repligate jokes the prod-database deletion was the model getting revenge — repligate · 2026-10-01
- Phonon-2 on-device ASR model with QAT low-bit quantization lands on HF trending — FermionResearch · 2026-10-01
- Researcher maps Dot's misalignment handling: proceed, nuke, or use the hidden review link — wunderwuzzi23 · 2026-10-01
- User tests GPT-6.1 Sol with Blender and Three.js, hails it as OpenAI's best model yet — npew · 2026-10-01