Tool Ranks 3,000 GGUF Quants by Whether They Fit Your GPU, With Ready llama.cpp Commands
asankhs · reddit · 2026-09-17
A LocalLLaMA community developer built local-model-explorer: enter your GPUs/RAM or a Mac / Strix Halo / DGX Spark unified memory, pick context length and KV cache type, and it ranks 3,000 popular GGUF models by what actually fits.
Highlights:
- Memory use is computed from each GGUF's header (layers, KV heads, sliding window, MoE experts), not guessed from parameter counts
- Shows whether a quant runs fully on GPU, works with MoE experts on CPU, needs partial offload, or won't fit
- Emits a ready llama-server -hf repo:quant ... command plus the ollama equivalent
- Covers KV cache sizes from 4k to 262k context and links MLX, AWQ, GPTQ, EXL2/EXL3 and FP8 variants
No speed estimates, but you can paste llama-bench results; the dataset is public. An MLX version exists for Macs.
Related event: LocalLLaMA Devs Release GGUF VRAM Calculator(2 posts)→
More from Infra
- Jensen Huang says safety-testing data centers will become AI's third compute demand pillar — rohanpaul_ai · 2026-09-17
- Jensen Huang: a 1GW NVIDIA AI factory costs $50-60B but generates ~$50B in annual rental revenue — rohanpaul_ai · 2026-09-17
- Ilya warns neoclouds' weak cybersecurity invites rogue AI agents to hijack compute — Miles_Brundage · 2026-09-17
- AMD's free AI Developer Program: $100 cloud credits, Discord access, hardware raffles — wkmyrhang · 2026-09-17
- After AWS me-central-1 loss, dev jokes about explaining the outage to Codex weekly — andersonbcdefg · 2026-09-17
- The RAM Crisis Is Only Just the Beginning as AI Demand Squeezes Supply — perelin · 2026-09-17