Real-world testing suggests Artificial Analysis Index is gamed and unrepresentative

PerformanceRound7913 · reddit · 2026-09-05

After hands-on testing, a Reddit user found Muse Spark 1.3 clearly underperforms Opus and SOL despite its high Artificial Analysis Index score, arguing the benchmark doesn't reflect real-world performance and is easy to game — raising questions about the trustworthiness of third-party LLM leaderboards.

Related event: User Tests Cast Doubt on Muse Spark 1.3 Benchmark Scores(4 posts)→

Original post →

More from Models

Models channel →