Together Offers Lowest Price and Highest Cache Hit Rate for Kimi K3 on OpenRouter
zhyncs42 · x · 2026-07-30
Data from the OpenRouter Kimi K3 dashboard reveals that Together Compute offers one of the lowest prices for Moonshot AI's Kimi K3 while achieving the highest prompt cache hit rate.
The combination of low pricing and high cache efficiency is critical for making long-context and agentic workloads more cost-effective. Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI with a 1M token context window.
More from Infra
- TensorSharp Adds Multi-GPU Tensor Parallelism, Boosting Local GGUF Inference Speeds — fuzhongkai · 2026-07-30
- NVIDIA Open-Sources PyCuTe: Pure Python Layout Algebra for CUTLASS — asdf1234_0 · 2026-07-30
- UC Berkeley's K-search: Auto-Translating CUDA Kernel Optimizations to Apple's MLX — berkeley_ai · 2026-07-30
- Advantech Edge Device Powered by Nvidia Thor Runs RealSense GMSL Cameras — chrismatthieu · 2026-07-30
- Samsung's Q2 Operating Profit Surges 1,800% to Record High Amid AI Chip Boom — Polymarket · 2026-07-30
- The Race for Power: Assessing Global Electricity Production for AI — lemire · 2026-07-30