Tool Ranks 3,000 GGUF Quants by Whether They Fit Your GPU, With Ready llama.cpp Commands

asankhs · reddit · 2026-09-17

A LocalLLaMA community developer built local-model-explorer: enter your GPUs/RAM or a Mac / Strix Halo / DGX Spark unified memory, pick context length and KV cache type, and it ranks 3,000 popular GGUF models by what actually fits.

Highlights:

No speed estimates, but you can paste llama-bench results; the dataset is public. An MLX version exists for Macs.

Related event: LocalLLaMA Devs Release GGUF VRAM Calculator(2 posts)→

Original post →

More from Infra

Infra channel →