Kimi K3 Review: The New Open-Source King with Stellar Search
andersonbcdefg · x · 2026-07-19
Independent reviewers added Kimi K3 to the prinzbench benchmark. K3 achieved an overall score of 47/99, far surpassing GLM-5.2's 30/99, making it the best open-source model tested to date.
- Overall Comparison: Its score is slightly below February's Gemini 3.1 Pro (50/99), putting them in the same tier. K3 excels slightly in search capabilities but is somewhat weaker in logical reasoning.
- Stability: K3 shows notable volatility; it occasionally misses simpler questions but frequently delivers stunning results on highly difficult ones.
- Core Highlight: With a search sub-score of 11/24, it becomes the first non-OpenAI model to correctly answer more than 10 search questions in this benchmark.
More from Models
- Claim says Kimi was distilled from Fable, sparking a model-attribution jab — cephaloform · 2026-07-22
- Gemini 3.6 Flash is now available in Antigravity and chat — MartianOnJupiter · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Macaron V1 adds LoRA RL on GLM 5.2 and claims SOTA benchmark gains — Xianbao_QIAN · 2026-07-22