Runescape Bench chart maps a crowded Pareto frontier across top models
teortaxesTex · x · 2026-07-23
A Runescape Bench chart shows a surprisingly diverse Pareto frontier across models, with several vendors occupying different points on the cost-performance curve.
The chart compares average API cost per run against benchmark performance and highlights a frontier that includes Fable 5, GPT-5.6 Terra xh, Grok 4.5 xh, GLM 5.2 fp4, DS V4 Flash, and Laguna S. The visual suggests there is no single dominant winner across all cost/performance tradeoffs.
More from Models
- Steve Hou expects a wave of U.S. open-source models as enterprise inference demand surges — soumitrashukla9 · 2026-07-23
- Musk says GPT-5 or GPT-6 could be indistinguishable from the smartest humans — kevinnbass · 2026-07-23
- One prompt was enough to get blocked, says an X user — gabriel1 · 2026-07-23
- Kimi K3 reportedly found and exploited a Redis 0day in 27 minutes with 32 agents — HanchungLee · 2026-07-23
- Reddit weighs a neglected MoE size class around 2B active parameters — WhoRoger · 2026-07-23
- ClinicalBench update shows Kimi K3 solving 7 of 10 EHR cases — teortaxesTex · 2026-07-23