PINNACLE Agentic Benchmark: Qwen3.8 Makes 15% Fewer Errors Than Claude Opus 5
ryanshrout · x · 2026-09-15
Signal65 scored eight new model configurations on its PINNACLE agentic benchmark, five of them open weights from Qwen, DeepSeek and Zai.
Key finding: Qwen3.8 made 15% fewer errors than Claude Opus 5 in their testing, and its 180B sibling runs on a desktop. The authors argue open weights are closer to the hosted frontier than the pacing debate admits — meaning any agreement to slow frontier development favors whoever can run their own hardware. Three labs atop the board (OpenAI, xAI, Anthropic) just backed a call to pace frontier AI.
Related event: PINNACLE Benchmark Update: Open Models Close Gap with Frontier(2 posts)→
More from Models
- Prefill and Decode: why asking an LLM for three takeaways from a long document still takes minutes — dotey · 2026-09-15
- Leak claims xAI trails OpenAI and Anthropic by roughly 6-12 months — sachinmaya1980 · 2026-09-15
- "They were way better last week": astra and fable regression claim — willcb · 2026-09-15
- Blogger: astra and fable are pacing the LLM frontier today — willcb · 2026-09-15
- User claims Claude Artifacts silently uploads drafts to cloud with toggle locked — maier_ak · 2026-09-15
- Kimi K3 is live and free on NVIDIA NIM with OpenAI-compatible API — airesearch12 · 2026-09-15