A BrowseComp chart puts Kimi K3 near the top on score while keeping cost low
iamfakhrealam · x · 2026-07-23
A chart compares several models on BrowseComp score vs. cost per task.
The visible points show:
- Kimi K3 (max) near the top-left, with a high BrowseComp score at low cost
- GPT-5.6 Sol and several Claude variants spread across higher-cost points
- Claude Mythos 5 (max) reaching the upper part of the chart at a relatively high cost
- Claude Opus 4.8 and Claude Sonnet 5 plotted below the top performers in score
The post’s main point is a pricing/performance comparison rather than a product announcement.
More from Models
- Gemma 4 tops 300 million downloads three months after launch — osanseviero · 2026-07-23
- OpenAI opens GPT-5.6 to the public across ChatGPT, Codex and API — emmanuelvivier · 2026-07-23
- Laguna at low quant seems to overthink and burn through context fast — IUseClifford · 2026-07-23
- Qwen-Image-3.0 gets put through layout-heavy tests against GPT-Image-2 — Scobleizer · 2026-07-23
- Dean Ball says Kimi is strong at coding, but open-weight economics still hurt — koltregaskes · 2026-07-23
- Simon Willison says loops are becoming obsolete as models handle long tasks natively — teropa · 2026-07-23