Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching
MatthewBerman · x · 2026-07-28
A post about Inference.net’s Catalyst Gateway describes a production workflow for evaluating Kimi K3 without switching traffic immediately.
The proposed rollout
- Install the Gateway.
- Keep sending production traffic to the current model.
- Let the gateway use an RLM to generate evals from live data for about 24 hours.
- Mirror live traffic to Kimi K3 while continuing to serve the old model.
- Once evals look good, get a Slack notification that it is safe to switch.
- Change the model ID to kimi-k3.
The pitch is that teams can cut token costs by around 50% while keeping control over their LLM stack.
More from Infra
- Database optimization drops CPU from 35% to 5% and p95 latency from 365 ms to 3 ms — DanielLockyer · 2026-07-28
- Local communities use Google and Meta data-center revenue to cut taxes and boost pay — aronchick · 2026-07-28
- Mapbox opens Places API with metadata for 250M places worldwide — anselm · 2026-07-28
- AI is likely to control quantum computers first, then use them for narrow science tasks — imjustnewatai · 2026-07-28
- Moonshot’s Kimi K3 lands in Japan with 2.8T open weights and $3/$13 pricing — DavidBennett__ · 2026-07-28
- A Kimi-k3 joke contrasts a $1,908 annual plan with $1.0908M to run it at home — HarveenChadha · 2026-07-28