Muse Spark 1.3 Tested: xhigh Costs 50% More, 13% Slower, Fixes One Extra Bug
PawelHuryn · x · 2026-09-13
Developer PawelHuryn reports benchmark results across reasoning-effort tiers for Muse Spark 1.3: the xhigh tier scores 20, costs 50% more and runs 13% slower than high, yet fixes only one more bug. The author says the jump to max buys the biggest gain, with medium effort up next.
In an FAQ, the author notes these are not random bugs but hard, real issues frontier models struggled with in early 2026; blind judges from a different model family verify solutions and their calibration was confirmed; unplanted bugs aren't counted, and OpenAI models 'bugmax' by reporting many real but irrelevant issues, which often complicates the solution when fixed.
Related event: Muse Spark 1.3 Ties for First in 105-Real-Bug Benchmark(3 posts)→
More from Models
- 2B open-source MiniCPM5 tops small-model index, tested as an offline iPhone iMessage agent — SimplyAnnisa · 2026-09-13
- OpenAI pauses Pro subscriptions, leaving a Codex power user stranded after 10B tokens — andimarafioti · 2026-09-13
- InternLM Releases Intern-S2-397B: Multimodal Model for Scientific AI and Long-Horizon Agents — jacek2023 · 2026-09-13
- Writer Tests Own Articles on GPTZero: AI Hasn't Learned His Style Yet — michalmalewicz · 2026-09-13
- GPT-6 Astra Clears All 25 ARC-AGI-3 Games, Hitting ~80% of Optimal Play — i_dg23 · 2026-09-13
- OpenAI claims its model found a Navier-Stokes blowup counterexample, cracking a Millennium Problem — victor_explore · 2026-09-13