One day with Kimi K3: high ceiling, high operational cost
Tarandjpop · reddit · 2026-07-21
## A full day with Kimi K3: high ceiling, heavy infrastructure demands The author says Kimi K3 can be the best model in the bunch if you let it think long enough, especially on visual taste and polish. But that strength comes with serious operational trade-offs. ### Main takeaways - **Reasoning is compulsory**: max reasoning cannot be turned off and consumes about **73–83%** of the output budget. - **Image tasks are slow**: some take **50–60 minutes** of reasoning, long enough to hit connection limits. - **Cheap on paper, expensive in practice**: the listed **$3/$15** pricing looks attractive, but real runs cost **2–3× Opus** because of reasoning-token usage. - **Infra burden**: it needs streaming, **100K–160K token** budgets, and hour-long connection lifetimes, forcing stack changes. - **Demand is overwhelming**: peak-time **429s** are common; only off-peak runs succeeded. The overall verdict is that K3’s quality ceiling is high, but its usability and infrastructure requirements are unusually punishing.
Related event: Kimi K3 Test Shows High Performance but 2-3x Cost of Claude Opus(2 posts)→
More from Infra
- AI performance is increasingly limited by materials science, not just compute — nordicinst · 2026-07-21
- A GLM-5.2 inference debate asks how 750B parameters can exceed 1 token per second — francoisfleuret · 2026-07-21
- Why vector databases slow AI agents down after constant writes — PrajwalTomar_ · 2026-07-21
- Larry Fink says China is ahead in the AI energy race, citing 100 GW nuclear buildout — rohanpaul_ai · 2026-07-21
- Local AI may pay back in 6–7 years and cut long-term costs by 30–40% — DavidLinthicum · 2026-07-21
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21