Fireworks says Kimi K3 matches Opus 5 closely on 663 coding tasks while costing 2.3x less
lqiao · x · 2026-07-28
Fireworks compares Kimi K3 with Opus 5 on 663 agentic coding tasks
- Fireworks says Kimi K3 and Opus 5 are very close in task-level quality across SWE (480), Algorithmic (100), and Terminal (83) tasks.
- On their benchmark, Opus 5 is slightly ahead overall on accuracy (92.6 vs 90.6 average), but Kimi K3 is much cheaper per task when priced via Fireworks serverless inference.
- Reported per-task costs:
- SWE: $0.52 vs $1.05 — Kimi 2.0x cheaper
- Algorithmic: $0.064 vs $0.18 — 2.8x cheaper
- Terminal: $0.35 vs $1.61 — 4.6x cheaper
- Average: $0.43 vs $0.99 — 2.3x cheaper
- The team says the key metric is per-task cost, not per-token cost, because open models often produce longer outputs.
- They say they will keep improving serving efficiency through Fireworks Inference and model quality through Fireworks Training.
Related event: Fireworks Tests Kimi K3: Quality Matches Opus 5 at 2-4.6x Lower Cost(5 posts)→
More from coding & agent
- Codex in Chat running for 12 hours is cited as a real agent example — ctjlewis · 2026-07-28
- A follow-up says the industry has been pushing “Agents™” since Q1 2024 — ctjlewis · 2026-07-28
- Vibe coding can turn everyday UI annoyances into a personal Chrome extension — tomchapin · 2026-07-28
- Salesforce adds Code Extension to Data 360 with Claude Code in the loop — msrivastav13 · 2026-07-28
- OpenAI agent demo, open-weight momentum and AI’s carbon cost reshape enterprise thinking — jonerp · 2026-07-28
- ChatGPT is already acting like an agent, but “Agents™” may be mostly branding — ctjlewis · 2026-07-28