smolbenchmark ranks sub-8GB models by speed, tok/J and heat on your own hardware
East-Muffin-6472 · reddit · 2026-09-12
- Mainstream leaderboards assume server GPUs; smolbenchmark targets models that fit in 8GB, ranked by decode speed, tokens per joule, and heat.
- Covers tablets, phones, Macs, Jetsons and Raspberry Pis with 13 model families so far and 1,000 configs on the Jetson nano Orin Super 8GB; a live device streams tok/s, tok/J, ITL, latency, power, thermals and battery.
- All raw data and reports are public; Pi, phone and Mac mini numbers are still being filled in, and feedback is welcome.
More from Infra
- Glass core substrates show 2x better warpage than organic core without stiffener — jwt0625 · 2026-09-12
- d-Matrix partners with NVIDIA to plug Raptor XPUs into NVLink Fusion rackscale systems — bookwormengr · 2026-09-12
- Running out of context on a large codebase: how to auto-handoff long-running local LLM tasks — Developer-Y · 2026-09-12
- The Economist: Nvidia is the central bank of AI — tolugenius · 2026-09-12
- What MoE/LLM runs well offline on a 24GB M5 MacBook Air? — itis_whatit-is · 2026-09-12
- Draft model hits ~60 tok/s running Qwen3.8-27B at 131k context on a 16GB GPU — pneuny · 2026-09-12