Blind doc benchmark puts Kimi K3 near GPT-5.6 Sol on Word files, but far behind on slides
ell-hol1 · reddit · 2026-07-21
Kimi K3 closes in on GPT-5.6 Sol in document writing, but not in slides
A Reddit-led blind benchmark called DocBench Arena collected 3,000+ votes on 236 generated files from 400+ users. On the combined leaderboard, GPT-5.6 Sol still ranks #1 with an ELO of 1306, while Kimi K3 debuts at #6 with 1192.
- Slides: Sol is far ahead at 1373 (#1) versus K3’s 1181 (#7).
- Documents: K3 nearly matches Sol, scoring 1214 (#5) vs 1211 (#6).
- Cost: K3 averages $0.44 per finished document, compared with $0.56 for Sol.
- Speed: Sol finishes a task in about 2 minutes, while K3 takes roughly 5× longer, uses more tokens, and needs more agent steps.
The author argues Sol still holds the crown because it dominates presentations and reaches a better efficiency frontier overall. K3’s main opening is Word documents, where it is already effectively tied on quality while being cheaper, though it still lags badly on speed. The sample size for K3 is also smaller, with 186 comparisons versus 423 for Sol, so the ranking could move as more votes come in.
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- giffmana: the env being used in training is part of the point — giffmana · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11