CAISI report says Kimi K3 leads open-weight models but trails U.S. frontier systems
mattsheehan88 · x · 2026-07-24
- A Commerce/CAISI-NIST report says the newest Chinese model is still behind top U.S. frontier models, but is the strongest open-weight model.
- The attached charts compare cyber capability and exploit-development performance across models, including Kimi K3 and GLM-5.2.
- In the exploit-development benchmark, top U.S. models score about 76.2%, while Kimi K3 is around 32.2% and GLM-5.2 around 24.4%.
- A second chart places Kimi K3 near the top of the Chinese model cluster on overall cyber capability, though still below the U.S. frontier band.
More from Models
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11