PINNACLE Benchmark Update: Open Models Close Gap with Frontier
Signal65's updated PINNACLE agentic benchmark tested 8 new model configs, 5 of them open-weight, finding Qwen3.8's error rate 15% lower than Claude Opus 5—evidence that open models rival frontier labs, undercutting calls to pause AI development.
2026-09-15 ~ 2026-09-15 · 2 related posts
- Signal65 tests show open-weight models closing on the frontier, undercutting pace calls — ryanshrout · 2026-09-15
- PINNACLE Agentic Benchmark: Qwen3.8 Makes 15% Fewer Errors Than Claude Opus 5 — ryanshrout · 2026-09-15