Anthropic benchmark score looks fine, but high ECI makes small gaps look bigger

scaling01 · x · 2026-07-25

The thread says the score itself looks fine: it is reportedly better than Opus 4.8, but still below Fable. The confusing part, they argue, is that Anthropic’s ECI is so high.

They also note that because ECI is sensitive to one-point differences, even small gaps can look more dramatic than they really are.

Related event: Anthropic's Benchmark Scores Spark Community Trust Crisis(2 posts)→

Original post →

More from Models

Models channel →