Muse Spark 1.3 looks benchmaxxed: real-world tests far below its benchmark scores
Swimming_Gain_4989 · reddit · 2026-09-05
A Reddit user tested Muse Spark 1.3 via both OpenRouter and Opencode Zen, with different system prompts and outside a harness, and found it consistently underperforming — suspecting the model is overfit to benchmarks. It was passable at Godot debugging and documentation lookups (roughly DeepSeek v4 flash level), but performed like GPT-4 mini at tutoring upper-level discrete math, unable to sustain back-and-forth conversation and often giving incomplete or incorrect explanations.
Related event: Hands-on tests call out Muse Spark 1.3 benchmark mismatch(4 posts)→
More from Models
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05