Kimi K3 impressed in a day of API tests, but reasoning tokens made real costs 2–3× Opus

Tarandjpop · reddit · 2026-07-21

A Reddit user spent a day testing **Kimi K3** against **Claude Opus** in production-like image-generation/editing workflows. Key observations: - **High output quality** when Kimi K3 finishes successfully - **Reasoning tokens dominated usage** at roughly **70–80%** of output tokens - Some end-to-end requests took **close to an hour** - Effective cost for this workload was often **2–3× higher** than Opus - Long-running requests forced infrastructure changes: **streaming, larger token budgets, longer-lived connections** - The main reliability issue was **peak-hour 429 / capacity errors** The poster says they were impressed by the ceiling, but the model currently has meaningful trade-offs in latency, infrastructure demands, and real-world cost for this workflow. The attached image compares Kimi K3, Opus, and GPT-5.6 Sol in a game-like benchmark-style visual.

Related event: Kimi K3 Test Shows High Performance but 2-3x Cost of Claude Opus(2 posts)→

Original post →

More from Models

Models channel →