PINNACLE Benchmark Update: Open Models Close Gap with Frontier

Signal65's updated PINNACLE agentic benchmark tested 8 new model configs, 5 of them open-weight, finding Qwen3.8's error rate 15% lower than Claude Opus 5—evidence that open models rival frontier labs, undercutting calls to pause AI development.

2026-09-15 ~ 2026-09-15 · 2 related posts