GPT-6.1 Sol Benchmarked Across All Effort Levels; Atra Faster but Pricier

PawelHuryn · x · 2026-10-01

Benchmarker PawelHuryn published results for GPT-6.1 Sol across all reasoning effort levels on his benchmark site, with n=3 runs for the max and xhigh levels and n=2 for the rest; additional n=3 runs are still in progress.

Compared with Astra: Astra is notably more expensive, but faster and spends fewer turns on the most challenging tasks. Live scores, caveats, and more models are available on the site.

Related event: GPT-6.1 Sol benchmarks land: near-Astra performance at a fraction of the cost(16 posts)→

Original post →

More from Models

Models channel →