From 100KB to 2.5TB: a developer maps the full AI model spectrum and what each tier costs to run

abhishekkumar333 · reddit · 2026-10-04

Arguing that 90% of AI use cases are over-engineered for datacenter-class models, the author maps the full model-size spectrum: 100KB TinyML models running on milliwatt microcontrollers, the 4GB–40GB local sweet spot (Mistral 7B, Gemma 2 9B/27B, Qwen 2.5 14B/32B at 4-bit on a Mac or RTX 3060/4090), and 2.5TB flagship MoE behemoths like DeepSeek and Kimi. They also published a deep dive on the VRAM math and a cheat sheet for matching model size to hardware.

Original post →

More from Infra

Infra channel →