A follow-up says Muse Spark beat Grok-4.20 in a private relationship benchmark

cis_female · x · 2026-07-21

The poster clarifies that the benchmark used cost numbers and model outputs to compare how different models commented on a relationship after ingesting roughly 600k tokens of private chat history.

The follow-up says people had treated Muse Spark and Grok-4.20 as roughly equivalent on price and quality, but this evaluation ranked Muse Spark near the top and Grok much worse.

Related event: Benchmark Compares LLMs on Cost-Efficiency and Deep Context Understanding(4 posts)→

Original post →

More from Models

Models channel →