Opus 5 looks far more cost-efficient than GPT-5.6 Sol on long-horizon agent tasks

daniel_mac8 · x · 2026-07-25

The post argues that critics misread the AA-Index cost-per-task chart when comparing Opus 5 to other models. Most AA-Index tasks are single-turn, so they do not capture the efficiency gains that appear on long-horizon agentic workloads.

The author points to ARC-AGI-3 as a better example: when a model can run for roughly 10,000 steps, they say Opus 5 becomes vastly more performant per unit cost than GPT-5.6 Sol.

Related event: Hyperagent Tests: Flagship Model Competition Shifts to Style and Cost(6 posts)→

Original post →

More from Models

Models channel →