Kimi K3 report adds in-house coding, agent, and WebDev benchmark tables

nrehiew_ · x · 2026-07-29

This follow-up adds two tables from the Kimi K3 report.

In-house benchmark table

It compares Kimi K3 (max) with several proprietary and open-weight models across coding, agent, and conversational benchmarks. Notable entries include:

WebDev bench

A second table compares Kimi K3 (max) against Claude Opus 4.8 (max) under blind expert judging on:

The report says experts judged outputs without knowing which model produced them, scoring code quality, feature completeness, visual fidelity, and interaction experience.

Related event: Deep Dive into Kimi K3 Tech Report: Architecture and Training(18 posts)→

Original post →

More from Infra

Infra channel →