Signal65 PINNACLE tests: Opus 5.5 cuts weighted errors 2.3x and price 34%; labs slash prices
ryanshrout · x · 2026-09-24
Signal65's PINNACLE enterprise benchmark data shows frontier labs halving prices this week:
- Claude Opus 5.5: 2.3x fewer weighted errors than Opus 5, 100% end-to-end multi-step completion (vs 95%), $0.40 per correct task (34% cheaper), list price 20% lower with cache reads at 5% of input (was 10%); retrieval fabrication dropped from 7.6% to 2.8%
- GPT-6 Sol: matches GPT-5.6 Sol at one-sixth the price
- GPT-6 Luna: cheapest correct task on the board at 1.5 cents, with the highest fabrication rate
Author's takeaway: capability isn't slowing down and neither are price cuts — you can only see both by measuring both.
More from Models
- TeleOCR, a Qwen2.5-VL-based document parsing model, trends on Hugging Face — StarDoc-AI · 2026-09-24
- Rumor: SSI to launch its first model this month after security-related delay — iruletheworldmo · 2026-09-24
- Which sub-40B finetunes work best for mimicking a writing style? — Borkato · 2026-09-24
- First-day Opus 5.5 verdict: power user says it replaced Astra entirely — kimmonismus · 2026-09-24
- Opus 5.5 impresses early users; Mirage launch video made with just 4 turns of edits — seanwbren · 2026-09-24
- Open source multimodal decision model XOR launches on Hugging Face, 260k context, Qwen-based — TheMoonMidas · 2026-09-24