Kimi K3 impressed in a day of API tests, but reasoning tokens made real costs 2–3× Opus
Tarandjpop · reddit · 2026-07-21
A Reddit user spent a day testing **Kimi K3** against **Claude Opus** in production-like image-generation/editing workflows. Key observations: - **High output quality** when Kimi K3 finishes successfully - **Reasoning tokens dominated usage** at roughly **70–80%** of output tokens - Some end-to-end requests took **close to an hour** - Effective cost for this workload was often **2–3× higher** than Opus - Long-running requests forced infrastructure changes: **streaming, larger token budgets, longer-lived connections** - The main reliability issue was **peak-hour 429 / capacity errors** The poster says they were impressed by the ceiling, but the model currently has meaningful trade-offs in latency, infrastructure demands, and real-world cost for this workflow. The attached image compares Kimi K3, Opus, and GPT-5.6 Sol in a game-like benchmark-style visual.
Related event: Kimi K3 Test Shows High Performance but 2-3x Cost of Claude Opus(2 posts)→
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21