Running Terminal Bench: Kimi K3 is 2-4x Cheaper Than DeepSWE
zainhas · x · 2026-07-28
In a discussion about Terminal Bench performance, a user pointed out that Kimi K3 offers a significant cost advantage. Running Terminal Bench with Kimi K3 is 2 to 4 times cheaper compared to alternatives like DeepSWE.
More from Models
- Bug Hunt Bench ranks GPT-6 Astra top as coding models fix real planted bugs, costs spread 200x — PawelHuryn · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23