Opus 5.5 at Medium Beats GPT-6 Astra Max on GDPval-AA at ~80% Lower Cost

nicolechirps · x · 2026-09-23

A comparison post claims that on the GDPval-AA benchmark, Claude Opus 5.5 at the medium setting outperforms GPT-6 Astra at its highest setting.

Estimated cost per task:

That's a better score at roughly 80% lower cost, per the post's screenshot. Benchmark details and sample sizes are not provided in the tweet itself.

Related event: Opus 5.5 mid-tier beats GPT-6 max at one-fifth the cost(3 posts)→

Original post →

More from Models

Models channel →