GGUF Fit Calculator Reads File Headers to Tell Which Quants Fit Your GPU
asankhs · reddit · 2026-09-17
A LocalLLaMA developer released local-model-explorer: enter your GPUs/RAM or unified-memory machines (Mac, Strix Halo, DGX Spark), pick context length and KV cache type, and it ranks 3,000 popular GGUF models by fit, with ready llama-server and ollama commands. Memory use is computed from GGUF headers rather than parameter-count guesses, distinguishing full-GPU, MoE-on-CPU, and partial-offload cases; an MLX version exists for Macs.
Related event: LocalLLaMA Devs Release GGUF VRAM Calculator(2 posts)→
More from Infra
- Jensen Huang: a 1GW NVIDIA AI factory costs $50-60B but generates ~$50B in annual rental revenue — rohanpaul_ai · 2026-09-17
- Ilya warns neoclouds' weak cybersecurity invites rogue AI agents to hijack compute — Miles_Brundage · 2026-09-17
- AMD's free AI Developer Program: $100 cloud credits, Discord access, hardware raffles — wkmyrhang · 2026-09-17
- After AWS me-central-1 loss, dev jokes about explaining the outage to Codex weekly — andersonbcdefg · 2026-09-17
- Crusoe runs 512 AMD MI355X GPUs at 5.75M tok/s in largest MLPerf inference entry — wkmyrhang · 2026-09-17
- Perovskite could lift solar efficiency ceiling from 30% to 45% — and give the US a shot against China — kyliebytes · 2026-09-17