From 100KB to 2.5TB: a developer maps the full AI model spectrum and what each tier costs to run
abhishekkumar333 · reddit · 2026-10-04
Arguing that 90% of AI use cases are over-engineered for datacenter-class models, the author maps the full model-size spectrum: 100KB TinyML models running on milliwatt microcontrollers, the 4GB–40GB local sweet spot (Mistral 7B, Gemma 2 9B/27B, Qwen 2.5 14B/32B at 4-bit on a Mac or RTX 3060/4090), and 2.5TB flagship MoE behemoths like DeepSeek and Kimi. They also published a deep dive on the VRAM math and a cheat sheet for matching model size to hardware.
More from Infra
- Lenovo's AI Express ships on-prem AI servers in 15 business days — shashib · 2026-10-04
- China's aggressive AI push runs into a new problem: too much usage, NYT reports — TMWNN · 2026-10-04
- UBS Lead Time Map: Foundries Take 3-4 Years While GPUs Ship in 6-12 Months — AccBalanced · 2026-10-04
- DeepSeek V4.1-Flash hits 5,800 TPS per Ascend 950DT card, but critics call throughput underwhelming — teortaxesTex · 2026-10-04
- NVIDIA NIM API unreliable despite decent speed, users seek alternatives — Professional_Log1367 · 2026-10-04
- LLMxRay: Open-Source Local Observability for LLM Traffic, Launches With One Command — GuruCsharp · 2026-10-04