GPT-5.6 Sol Matches Claude in Economic Task Benchmark

thesaraharminta · x · 2026-07-10

Reports indicate that GPT-5.6 Sol (max) achieved scores very close to Claude Fable 5 (max) in the GDPval-AA v2 benchmark.

The post emphasizes that this result reflects a roughly equivalent level of capability between the two models in completing economic value tasks.

Original post →

More from Models

Models channel →