New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi
maximelabonne · x · 2026-09-04
Maxime Labonne shared the Localmaxxing speed benchmark, which tracks decode/prefill tok/s, TTFT and VRAM for small models like LFM2-350M-NVFP4A16 across hardware from RTX 5090 down to Raspberry Pi, plus a decode calculator and community submissions — handy for picking edge deployment setups.
More from Infra
- Perplexity CEO pitches local AI hardware with Nvidia and Apple, backs hybrid cloud-local — Kr00ney · 2026-09-05
- Apple vs. NVIDIA: the balance sheet as competitive advantage, B2C then, B2B2B now — BenBajarin · 2026-09-05
- Anima's Accelerated Understanding Bets on a Foundation Model for the Physical World — Latent Space · 2026-09-05
- Dev forks vLLM with custom patch to benchmark 31B model unsupported by flashinfer — abhijithneil · 2026-09-04
- Investor Gavin Baker: AI data centers are 'the best thing' for US working class — GavinSBaker · 2026-09-04
- Kafka, Kafka Connect and Schema Registry exposed as native MCP tools — jkriket · 2026-09-04