A model leaderboard shows Fable scores stay stable while rankings barely move

cis_female · x · 2026-07-22

A ranking table shows Fable-style scores stay mostly stable across runs

The post says there is some grading variability in Fable, but not much, and the rankings remain mostly the same.

The attached chart shows a model leaderboard with mean score, standard deviation, and letter grade. It ranks models such as anthropic/claude-opus-4.8, meta/muse-spark-1.1, openai/gpt-5.6-sol, z-ai/glm-5.2, google/gemini-3.5-flash, and x-ai/grok-4.20.

The reply adds a cost-based angle: people had been treating muse spark and grok-4.20 as roughly equivalent cheap-good models, but this table suggests muse spark is near the top while grok is far lower.

Related event: Testing LLMs on Relationship Understanding with 600k Tokens(6 posts)→

Original post →

More from Models

Models channel →