Kimi K3 Lags Behind Muse Spark in Evaluation

adonis_singh · x · 2026-07-17

In the eyebench test, **muse-spark-1.1** performed exceptionally well, essentially outperforming other models except for gpt-5.4/5.6. The author pointed out that **kimi-k3** scored about 4 points lower than muse, and its testing cost was roughly 10 times higher over 100 questions, making the results "not look great."

Related event: Kimi K3 Lags Behind Muse Spark in Benchmarks with Higher Costs(2 posts)→

Original post →

More from Models

Models channel →