Open-source zxLLM predicts LLM VRAM usage & KV-cache needs with high precision

Capable_Item_5918 · reddit · 2026-08-17

A developer released zxLLM, an open-source tool designed to accurately predict VRAM usage and KV-cache requirements for local LLMs. It supports vLLM, SGLang, and llama.cpp, automatically detecting architectural nuances like GQA and querying local GPU states via nvidia-smi. Tested on models like Llama 3, Qwen 2.5, and DeepSeek, it achieves a low error rate (0.04% - 1.8%) and runs with zero external dependencies.

Original post →

More from coding & agent

coding & agent channel →