LFM2-350M NVFP4A16 hits 1.7M tok/s decode on a single RTX 5090
AlpinDale · x · 2026-09-04
The Localmaxxing leaderboard lists Liquid AI's LFM2-350M quantized to NVFP4A16, reaching roughly 1.7M tok/s decode on a single RTX 5090 (32 GB). AlpinDale quipped that at that speed you "might as well read /dev/urandom." The tweet is just the quip plus a link; detailed latency and VRAM numbers live on the linked leaderboard page.
More from Infra
- AMD's Threadripper Halo Station packs 96 cores and 576GB of HBM3E — ccerrato147 · 2026-09-04
- AeroJEPA fluid foundation model joins NVIDIA's PhysicsNeMo ecosystem — ricardovinuesa · 2026-09-04
- Building a €2-2.5k local AI rig for legal RAG and agentic coding: hardware picks debated — whatyathinkk · 2026-09-04
- Dual 3090 owners debate adding more cards: bigger local models vs parallel instances — Blues520 · 2026-09-04
- NousResearch brings one-click local model setup to Hermes Agent on NVIDIA systems — lifebypixels · 2026-09-04
- Regulated-industry dev seeks AI Gateway with Okta SSO and runtime policy enforcement — IrrepressibleInk · 2026-09-04