Kimi K3 report adds in-house coding, agent, and WebDev benchmark tables
nrehiew_ · x · 2026-07-29
This follow-up adds two tables from the Kimi K3 report.
In-house benchmark table
It compares Kimi K3 (max) with several proprietary and open-weight models across coding, agent, and conversational benchmarks. Notable entries include:
- Coding Experience: Kimi K3 scores competitively across Claude Code and Codex harnesses.
- General Agent Experience: strong results on benchmarks such as 24/7 ClawBench 2.0, MIRA Bench, KAET, CLIF Bench, Agentic Vision Bench, Swarm Bench, Online Experience, Deep Research Bench, Finance Bench, KWV Bench, DECK Bench, and Agent Behavior Bench.
- Conversational Experience: Kimi K3 performs strongly on Faithfulness and Chat All-in-One Bench.
WebDev bench
A second table compares Kimi K3 (max) against Claude Opus 4.8 (max) under blind expert judging on:
- Games
- 3D / WebGL / Shader
- Website / UI Clone
- Overall
The report says experts judged outputs without knowing which model produced them, scoring code quality, feature completeness, visual fidelity, and interaction experience.
Related event: Deep Dive into Kimi K3 Tech Report: Architecture and Training(18 posts)→
More from Infra
- Kernel Forge uses MCTS to optimize CUDA kernels and beats PyTorch baselines on 14 cases — omarsar0 · 2026-07-29
- X Spaces AMA says the AI race has shifted from models to compute and energy — AIFlow_ML · 2026-07-29
- AI race has moved past models and into compute, energy, and infrastructure — AIFlow_ML · 2026-07-29
- Google Cloud Run sandbox shows why isolating untrusted code still matters — rseroter · 2026-07-29
- Structural Shift in Semi Supply Chain: Vendors Secure Rare Long-Term Customer Commitments — BenBajarin · 2026-07-29
- Lightning AI says July brought new infra, a bigger Lightning Cloud, and better performance — LightningAI · 2026-07-29