A follow-up says Muse Spark beat Grok-4.20 in a private relationship benchmark
cis_female · x · 2026-07-21
The poster clarifies that the benchmark used cost numbers and model outputs to compare how different models commented on a relationship after ingesting roughly 600k tokens of private chat history.
The follow-up says people had treated Muse Spark and Grok-4.20 as roughly equivalent on price and quality, but this evaluation ranked Muse Spark near the top and Grok much worse.
Related event: Benchmark Compares LLMs on Cost-Efficiency and Deep Context Understanding(4 posts)→
More from Models
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- pi 0.81.0 adds first-class integration with llama.cpp server — huggingface · 2026-07-21
- GPT-5.6 Sol is said to explain weak opinions better than Opus 4.8 — eyishazyer · 2026-07-21
- Leak claims Gemini 3.5 Pro gets a 2M-token context and better coding — bdsqlsz · 2026-07-21
- Ethan Mollick says AI-writing sameness is about craft, not em dashes or bullet lists — emollick · 2026-07-21
- Kimi rolls out paid plan upgrades with tiers from $19 to $199 a month — gnukeith · 2026-07-21