Signal65 tests show open-weight models closing on the frontier, undercutting pace calls
ryanshrout · x · 2026-09-15
As OpenAI, xAI and Anthropic backed a call to pace frontier AI development, Signal65's PINNACLE agentic benchmark scored eight new configurations — five of them open-weight from Qwen, DeepSeek and Zai.
Key results:
- Open-weight Qwen3.8-2.4T-A95B made fewer weighted errors than Claude Opus 5
- DeepSeek-V4.1-Flash cut errors by a third vs. its previous generation at max effort
- Gemini 3.8 Flash completed 99.6% of multi-step jobs
Argument: "You cannot pace a frontier you have not measured." PINNACLE scores real multi-step enterprise work with code-verified, deterministic scoring (no model judges), measuring correct work, speed and cost. Open weights are closer to the hosted frontier than the pacing conversation admits, leaving little room to slow down. NVIDIA (Blackwell Ultra as reference platform) and AMD both endorsed the benchmark.
Related event: PINNACLE Benchmark Update: Open Models Close Gap with Frontier(2 posts)→
More from Models
- User verdict after testing small models: DeepSeek v4.1 Flash is the minimum viable model for real work — solyarisoftware · 2026-09-15
- Raschka deep-dive: GPT-6 Astra, looped transformers, and the hidden chain-of-thought question — AxSaucedo · 2026-09-15
- Rumors: Anthropic quietly routing traffic to an unannounced Opus 5.2 — kimmonismus · 2026-09-15
- OpenAI: new model solved Navier–Stokes in 88 hours with 10,000 coordinating agents — ZeroStateReflex · 2026-09-15
- OpenAI researcher explains why lab staff are suddenly scared about AI progress — socoolandawesome · 2026-09-15
- Local Qwen loops and forgets in coding agents while Claude Code just works — tlpta · 2026-09-15