PINNACLE Agentic Benchmark: Qwen3.8 Makes 15% Fewer Errors Than Claude Opus 5

ryanshrout · x · 2026-09-15

Signal65 scored eight new model configurations on its PINNACLE agentic benchmark, five of them open weights from Qwen, DeepSeek and Zai.

Key finding: Qwen3.8 made 15% fewer errors than Claude Opus 5 in their testing, and its 180B sibling runs on a desktop. The authors argue open weights are closer to the hosted frontier than the pacing debate admits — meaning any agreement to slow frontier development favors whoever can run their own hardware. Three labs atop the board (OpenAI, xAI, Anthropic) just backed a call to pace frontier AI.

Related event: PINNACLE Benchmark Update: Open Models Close Gap with Frontier(2 posts)→

Original post →

More from Models

Models channel →