A model benchmark shows Muse Spark far ahead of Grok-4.20 on score vs cost
cis_female · x · 2026-07-21
The attached comparison table ranks several models by score and first-call cost.
Top of the table is Claude Fable 5 at 4.30, followed by Claude Opus 4.8 at 4.06. Meta’s Muse Spark 1.1 scores 3.88 at a much lower listed cost, while GLM-5.2, GPT-5.6 variants, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Plus, Grok-4.20, and MiniMax M3 trail behind.
The poster’s point is that the cheap-new-model narrative does not hold uniformly: some low-cost models perform very well, but others do not.
Related event: Testing LLMs on Relationship Understanding with 600k Tokens(6 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11