KV Cache Size Calculator
zephyr_z9 · x · 2026-07-18
Introduces a KV Cache Size Calculator tool designed to estimate the cache footprint of various models based on token length, sequence count, KV precision, and indexer precision.
Examples for DeepSeek V4 Pro and GLM-5.2 are provided: under 1 million tokens, single sequence, FP8/INT8 KV precision, and FP4/INT4 indexer precision, their total cache sizes are approximately 4.14 GiB and 43.09 GiB, respectively. The page also details the formulas and breakdowns, showing its capability to estimate inference KV/indexer costs across different architectures.
More from Infra
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21