Kimi K3 impressed in a day of API tests, but reasoning tokens made real costs 2–3× Opus
Tarandjpop · reddit · 2026-07-21
A Reddit user spent a day testing Kimi K3 against Claude Opus in production-like image-generation/editing workflows.
Key observations:
- High output quality when Kimi K3 finishes successfully
- Reasoning tokens dominated usage at roughly 70–80% of output tokens
- Some end-to-end requests took close to an hour
- Effective cost for this workload was often 2–3× higher than Opus
- Long-running requests forced infrastructure changes: streaming, larger token budgets, longer-lived connections
- The main reliability issue was peak-hour 429 / capacity errors
The poster says they were impressed by the ceiling, but the model currently has meaningful trade-offs in latency, infrastructure demands, and real-world cost for this workflow. The attached image compares Kimi K3, Opus, and GPT-5.6 Sol in a game-like benchmark-style visual.
Related event: Kimi K3 Test Shows High Performance but 2-3x Cost of Claude Opus(2 posts)→
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11