DeepSWE Eval: Kimi K3 Matches Claude Fable 5 at 35% of the Cost
togethercompute · x · 2026-07-22
Together AI conducted a model comparison for software engineering tasks based on the DeepSWE benchmark.
Results show that Kimi K3 achieves the same performance level as Claude Fable 5 at approximately 35% of the price. Furthermore, at higher pass@k settings (allowing more generation attempts), Kimi K3 actually pulls ahead of the competition.
Related event: Kimi K3 Matches Fable 5 in SWE Tasks at a Third of the Cost(7 posts)→
More from Models
- Open-Weight Model Hy3 Ranks #16 in Frontend Code Arena — arena · 2026-07-22
- Gemini 3.5 Flash Outperforms GPT-5.6 in Light Coding Tasks — Shick_hydro · 2026-07-22
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22