Government testing reportedly finds Kimi K3 far behind U.S. frontier models
Polymarket · x · 2026-07-24
Polymarket says US and UK government testing found Moonshot’s Kimi K3 performed “significantly below” US frontier models.
This is not an official benchmark release, but it is a notable performance claim about a Chinese frontier model being tested against leading US systems.
More from Models
- Bindu Reddy says Kimi K3 is cheap, strong on long tasks, but not frontier — bindureddy · 2026-07-24
- GPT-5.6 Pro beats Codex on critique and repair, says Will Depue — willdepue · 2026-07-24
- Artificial Analysis billboards rank leading models by intelligence and cost per task — ArtificialAnlys · 2026-07-24
- Grok 4.5 is now available to all accounts across web, X, iOS, and Android — Ready-Independent108 · 2026-07-24
- Laguna-S-2.1 infinite-thinking loops may come from quantization, not prompting — CautiousStudent6919 · 2026-07-24
- User Cancels Claude Max After Weeks of Talking, Saying the Model Is Just Too Moralistic — breath_mirror · 2026-07-24