Kimi-K3 Performance on FrontierMath Benchmark

scaling01 · x · 2026-07-19

Moonshot's Kimi-K3 (max) model scored 39% on the FrontierMath Tier 4 benchmark.

This score is 7% lower than the best US model from 7 months ago.

Related event: Kimi K3 Evaluations Show Polarized Results and Harness Sensitivity(5 posts)→

Original post →

More from Models

Models channel →