Signal65 tests show open-weight models closing on the frontier, undercutting pace calls

ryanshrout · x · 2026-09-15

As OpenAI, xAI and Anthropic backed a call to pace frontier AI development, Signal65's PINNACLE agentic benchmark scored eight new configurations — five of them open-weight from Qwen, DeepSeek and Zai.

Key results:

Argument: "You cannot pace a frontier you have not measured." PINNACLE scores real multi-step enterprise work with code-verified, deterministic scoring (no model judges), measuring correct work, speed and cost. Open weights are closer to the hosted frontier than the pacing conversation admits, leaving little room to slow down. NVIDIA (Blackwell Ultra as reference platform) and AMD both endorsed the benchmark.

Related event: PINNACLE Benchmark Update: Open Models Close Gap with Frontier(2 posts)→

Original post →

More from Models

Models channel →