Kimi K3 is the First Open-Source Model to Pass Compound's Internal Benchmark

peterjliu · x · 2026-07-29

The Compound team shared recent test results from their multi-model orchestration system, which leverages orchestrator and sub-agents to combine the strengths of various models, achieving Pareto points unattainable by any single lab or model.

In continuous testing against their internal 'checklist' benchmark, only recent Claude models consistently passed, Gemini generally failed, and OpenAI only recently passed with Sol 5.6. The latest testing reveals that Moonshot's Kimi K3 is the first open-source model to successfully pass this benchmark.

Related event: Kimi K3 Passes Compound Benchmark Amid Cost Efficiency Concerns(2 posts)→

Original post →

More from Models

Models channel →