Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks
ChrisGPT · x · 2026-07-22
Kimi K3 beats Gemini 3.6 Flash on four shared benchmarks
The post claims that Kimi K3, described as an open-source model, outperformed Google’s Gemini 3.6 Flash on every shared benchmark shown in the chart.
- DeepSWE v1.1: 67.5% vs 49.0% (+18.5)
- Terminal-Bench 2.1: 88.3% vs 78.0% (+10.3)
- GDPVal-AA v2: 1668 vs 1421 (+247)
- CharXiv Reasoning: 91.3% vs 89.4% (+1.9)
The post frames Kimi K3 as winning across coding, agentic terminal work, knowledge tasks, and chart reasoning.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11