Kimi K3 tops Gemini 3.6 Flash on four shared public benchmarks
ChrisGPT · x · 2026-07-22
Kimi K3 beats Gemini 3.6 Flash on four shared benchmarks
The post claims that Kimi K3, described as an open-source model, outperformed Google’s Gemini 3.6 Flash on every shared benchmark shown in the chart.
- DeepSWE v1.1: 67.5% vs 49.0% (+18.5)
- Terminal-Bench 2.1: 88.3% vs 78.0% (+10.3)
- GDPVal-AA v2: 1668 vs 1421 (+247)
- CharXiv Reasoning: 91.3% vs 89.4% (+1.9)
The post frames Kimi K3 as winning across coding, agentic terminal work, knowledge tasks, and chart reasoning.
More from Models
- Users are switching GPT-5.6 variants to dodge cybersecurity request blocks — ivan_bezdomny · 2026-07-22
- A discussion of post-training incentives and long-horizon instruction following — xuanalogue · 2026-07-22
- A frontier AI control stack proposes logs, scans, defenses, and breach plans — sjgadler · 2026-07-22
- Fireworks says Kimi K3 handles 72–96% of agent traffic at up to 50x lower cost — eliebakouch · 2026-07-22
- Mako debuts as a web-native model for end-to-end live web execution — Scobleizer · 2026-07-22
- GPT-6 is said to be near, with OpenAI betting on faster inference and custom chips — haider1 · 2026-07-22