Muse Spark 1.3 Tested: xhigh Costs 50% More, 13% Slower, Fixes One Extra Bug

PawelHuryn · x · 2026-09-13

Developer PawelHuryn reports benchmark results across reasoning-effort tiers for Muse Spark 1.3: the xhigh tier scores 20, costs 50% more and runs 13% slower than high, yet fixes only one more bug. The author says the jump to max buys the biggest gain, with medium effort up next.

In an FAQ, the author notes these are not random bugs but hard, real issues frontier models struggled with in early 2026; blind judges from a different model family verify solutions and their calibration was confirmed; unplanted bugs aren't counted, and OpenAI models 'bugmax' by reporting many real but irrelevant issues, which often complicates the solution when fixed.

Related event: Muse Spark 1.3 Ties for First in 105-Real-Bug Benchmark(3 posts)→

Original post →

More from Models

Models channel →