Repost accuses Claude benchmark chart of highlighting only GPT-5.6 Sol’s sole win

soumitrashukla9 · x · 2026-07-25

A repost accuses Claude’s benchmark graphic of being misleading because it highlights only the one benchmark where GPT-5.6 Sol won.

The post quotes Anthropic’s claim that Opus 5 is the new state of the art on several coding and knowledge-work evaluations, and the commenter argues the chart presentation is “so dishonest.”

Original post →

More from Models

Models channel →