Meta's Muse Spark 1.3: 38% fewer errors than Opus 5 at $0.85 per correct task
ryanshrout · x · 2026-09-04
Signal65's Ryan Shrout published new PINNACLE agentic benchmark results: Meta's Muse Spark 1.3 ranks second, with 38% fewer weighted errors than Claude Opus 5, and costs $0.85 per correct task versus $1.36 for Opus and $1.22 for GPT-5.6 Sol — knocking both hosted frontier models off the cost frontier.
Details:
- The model is slow and token-hungry, yet still the best buy at the top of the hosted tier
- PINNACLE builds real multi-step enterprise tasks from US Department of Labor work activities, with deterministic code-verified scoring, no model judges, and a freshly generated sandbox and answer key each run to prevent memorization
- Full results are available on the Signal65 site
This may be Meta's next "Llama moment."
More from Models
- GPT-6 Astra beats Fable 5.1 on Terminal science and Automation benchmarks — ChrisGPT · 2026-09-04
- Unconfirmed GPT-6 Astra benchmarks leak amid 'AGI era' launch claims — mark_k · 2026-09-04
- GPT-6 Astra priced at $10/1M input, $50/1M output, same as Fable 5.1 — bindureddy · 2026-09-04
- Astra Priced at $10/M Input and $50/M Output Tokens — saln1 · 2026-09-04
- GPT-6 Astra vs Gemini 3.8 Flash: user posts head-to-head comparison — Able-Line2683 · 2026-09-04
- GPT-6 Astra pricing leaked: $10/M input, $50/M output tokens, near-perfect ARC-AGI-3 — thesaraharminta · 2026-09-04