Kimi K3 Lags Behind Muse Spark in Evaluation

adonis_singh · x · 2026-07-17

In the eyebench test, muse-spark-1.1 performed exceptionally well, essentially outperforming other models except for gpt-5.4/5.6.

The author pointed out that kimi-k3 scored about 4 points lower than muse, and its testing cost was roughly 10 times higher over 100 questions, making the results "not look great."

Related event: Kimi K3 Lags Behind Muse Spark in Benchmarks with Higher Costs(2 posts)→

Original post →

More from Models

Models channel →