Private benchmark says Kimi K3 beats GPT 5.6 Sol on long-horizon office work
ZainHasan6 · x · 2026-07-22
- A private evaluation from AA says Kimi K3 beats GPT 5.6 Sol on a real-world long-horizon knowledge-work benchmark.
- The same evaluation places Kimi K3 second only to Fable 5.
- The benchmark focuses on deliverables such as spreadsheets, presentations, and memos, not just abstract model scores.
Related event: Kimi K3 Beats GPT-4o in Long-Horizon Task Benchmark(2 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11