New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi

maximelabonne · x · 2026-09-04

Maxime Labonne shared the Localmaxxing speed benchmark, which tracks decode/prefill tok/s, TTFT and VRAM for small models like LFM2-350M-NVFP4A16 across hardware from RTX 5090 down to Raspberry Pi, plus a decode calculator and community submissions — handy for picking edge deployment setups.

Original post →

More from Infra

Infra channel →