smolperfbenchmark: a public leaderboard for small open models on 8GB-class devices
East-Muffin-6472 · reddit · 2026-09-16
A new open benchmark, smolperfbenchmark, ranks small open-source models on consumer hardware instead of assuming server GPUs. All models run the same GGUFs, comparing llama.cpp vs Ollama with locked power modes and logged thermals.
First data point: SmolLM2-135M hits 165 tok/s and 29.6 tokens/joule at 25W on an 8GB Jetson Orin Nano Super. The live board covers 13 model families with 1000 configurations on that device, measuring tok/s, tok/J, ITL, latency, power, thermals, and battery. Target hardware ranges from tablets and phones to Macs, Jetsons, and Raspberry Pis, with Pi/phone/Mac mini data still being filled in. All results will be publicly released so users can pick the right model for their own hardware.
More from Infra
- Nvidia says DSX fits 40% more GPUs per watt; Vera Rubin to cut token cost ~45x — VishalG · 2026-09-16
- Nomura sees memory market hitting $3.68T by 2030, 83% from data centers — Beth_Kindig · 2026-09-16
- SOTA GGUF quants for Qwen3.8-Flash-Next: near-baseline quality at 1/5 the size — BullfrogScary8947 · 2026-09-16
- Baidu Cloud pitches 'industrial agents': 80% of SOEs onboard, AI revenue hits $1.75B — 智东西 · 2026-09-16
- Huawei KunLan launches 20 products, AI agent cuts network setup from 1 hour to 5 minutes — 智东西 · 2026-09-16
- Ministral 3 3B on a Galaxy S21 relays chats between Gemini and Z.ai across two browsers, 10/10 runs — Mean-Standard7390 · 2026-09-16