Why a better benchmark score may still feel like a weaker model in practice
robleclerc · x · 2026-07-28
The quoted thread argues that benchmark rankings can diverge from real-world model quality: a smaller or more efficiently trained model may look better on tests, while a larger one may still feel smarter in use.
The main explanation offered is that RL on bigger models is harder and more expensive. As a result:
- Smaller models can absorb new data, environments, and RL compute faster.
- Larger models may improve more slowly because each RL cycle is costlier.
- That could let a model like Opus keep winning benchmarks even if another model feels stronger in practice.
The reply adds a broader human parallel: some people are very good at extracting and repackaging knowledge from others even when they are not the deepest thinkers themselves.
Related event: AI Benchmark Scores Disconnect from Real-World Usage(3 posts)→
More from Models
- Kimi K3 throughput jumps from 19 to 49 tok/s on OpenRouter — cedric_chee · 2026-07-28
- Kimi K3 launches with 2.8T parameters, 1M context and $0.30 input pricing — togethercompute · 2026-07-28
- Together AI brings Moonshot’s Kimi K3 online with 1M context and agent tools — togethercompute · 2026-07-28
- k3 adds no RL algorithm changes, while tool-call steps track eval scores almost 1:1 — stochasticchasm · 2026-07-28
- DeepInfra adds Claude Opus 5 with 1M context and $5/$25 pricing — gharik · 2026-07-28
- Kimi K3 goes live on OpenRouter as third-party providers race to add support — scaling01 · 2026-07-28