Artificial Analysis launches Endpoint Accuracy Index: GLM-5.2, gpt-oss-120b, DeepSeek V4 Pro show big accuracy gaps across API providers

ArtificialAnlys · x · 2026-08-05

Artificial Analysis has launched the Endpoint Accuracy Index, measuring how much of an open-weights model's accuracy each serverless API endpoint preserves. Initial coverage includes GLM-5.2, gpt-oss-120b, and DeepSeek V4 Pro, with Kimi K3 coming soon.

The index benchmarks each endpoint against a self-hosted reference deployment of official weights (100%), using three equally weighted areas: tool calling (BFCL-500), scientific reasoning (HLE-250), and long-context recall (AA-LCR-25), with confidence intervals.

Key findings:

Related event: Artificial Analysis Launches Endpoint Accuracy Index(4 posts)→

Original post →

More from Infra

Infra channel →