Same Model, Different Quality: Endpoint Accuracy Varies 73%–100% Across Providers

ArtificialAnlys · x · 2026-08-21

Artificial Analysis hosted its latest "Inference, Measured" event in San Francisco, exploring how the same model is not always the same product — covering serverless inference, the Endpoint Accuracy Index, and AA-AgentPerf.

The new Endpoint Accuracy Index works by self-hosting released weights as a 100% reference, then measuring each provider's endpoint with the same three evaluations and scoring relative to that reference. Results ranged from 73% to 100% across providers.

Quantization, KV-cache compression, and context limits can all degrade quality — differences that never show up on a pricing page. Full results are available on their website.

Original post →

More from Infra

Infra channel →