A model leaderboard shows Fable scores stay stable while rankings barely move
cis_female · x · 2026-07-22
A ranking table shows Fable-style scores stay mostly stable across runs
The post says there is some grading variability in Fable, but not much, and the rankings remain mostly the same.
The attached chart shows a model leaderboard with mean score, standard deviation, and letter grade. It ranks models such as anthropic/claude-opus-4.8, meta/muse-spark-1.1, openai/gpt-5.6-sol, z-ai/glm-5.2, google/gemini-3.5-flash, and x-ai/grok-4.20.
The reply adds a cost-based angle: people had been treating muse spark and grok-4.20 as roughly equivalent cheap-good models, but this table suggests muse spark is near the top while grok is far lower.
Related event: Testing LLMs on Relationship Understanding with 600k Tokens(6 posts)→
More from Models
- Is DeepSeek's rumored K3 a scaled-down model, or something bigger? X users debate — teortaxesTex · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11