GGUF Fit Calculator Reads File Headers to Tell Which Quants Fit Your GPU

asankhs · reddit · 2026-09-17

A LocalLLaMA developer released local-model-explorer: enter your GPUs/RAM or unified-memory machines (Mac, Strix Halo, DGX Spark), pick context length and KV cache type, and it ranks 3,000 popular GGUF models by fit, with ready llama-server and ollama commands. Memory use is computed from GGUF headers rather than parameter-count guesses, distinguishing full-GPU, MoE-on-CPU, and partial-offload cases; an MLX version exists for Macs.

Related event: LocalLLaMA Devs Release GGUF VRAM Calculator(2 posts)→

Original post →

More from Infra

Infra channel →