Artificial Analysis launches Endpoint Accuracy Index: GLM-5.2, gpt-oss-120b, DeepSeek V4 Pro show big accuracy gaps across API providers
ArtificialAnlys · x · 2026-08-05
Artificial Analysis has launched the Endpoint Accuracy Index, measuring how much of an open-weights model's accuracy each serverless API endpoint preserves. Initial coverage includes GLM-5.2, gpt-oss-120b, and DeepSeek V4 Pro, with Kimi K3 coming soon.
The index benchmarks each endpoint against a self-hosted reference deployment of official weights (100%), using three equally weighted areas: tool calling (BFCL-500), scientific reasoning (HLE-250), and long-context recall (AA-LCR-25), with confidence intervals.
Key findings:
- GLM-5.2: Output token limits severely restrict accuracy; the most restrictive endpoints score half the reference or less on HLE-250.
- gpt-oss-120b: Tool call handling varies widely; some endpoints score 22% on BFCL-500 vs 37% for the reference. Serving configuration changes model behavior, and restricted context windows truncate long tasks.
- DeepSeek V4 Pro: Most endpoints are at reference parity, with DeepSeek's own first-party endpoint performing well.
Related event: Artificial Analysis Launches Endpoint Accuracy Index(4 posts)→
More from Infra
- Burning $130K/Day? Unpacking DeepSeek API Token Volumes — teortaxesTex · 2026-08-05
- NSF Launches $100M Program for Regional AI Infrastructure Hubs — mkratsios47 · 2026-08-05
- Chutes AI Enforces TEE Verification: 8x RTX 5090s Beat Pro GPUs at 65% Lower Cost — markjeffrey · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05
- LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context — helloiamleonie · 2026-08-05
- ai& Partners with Voltaiq for Battery Storage in Japanese AI Data Centers — DavidBennett__ · 2026-08-05