Muse Spark 1.3 looks benchmaxxed: real-world tests far below its benchmark scores

Swimming_Gain_4989 · reddit · 2026-09-05

A Reddit user tested Muse Spark 1.3 via both OpenRouter and Opencode Zen, with different system prompts and outside a harness, and found it consistently underperforming — suspecting the model is overfit to benchmarks. It was passable at Godot debugging and documentation lookups (roughly DeepSeek v4 flash level), but performed like GPT-4 mini at tutoring upper-level discrete math, unable to sustain back-and-forth conversation and often giving incomplete or incorrect explanations.

Related event: Hands-on tests call out Muse Spark 1.3 benchmark mismatch(4 posts)→

Original post →

More from Models

Models channel →