Local LLM quantization guide: Hardware thresholds for FP8, NVFP4, and more

Ill_Dragonfruit_3547 · reddit · 2026-08-20

The author shares lessons on quantization and hardware compatibility for running local LLMs, emphasizing that quant formats are tied to specific hardware generations, not just compression ratios.

Hardware Generations & Formats:

Format Taxonomy:

Key Rule: On an M1Max, stick to MLX for LLMs and GGUF/fp16 safetensors for diffusion. Other formats are likely to fail or perform poorly.

Original post →

More from Infra

Infra channel →