LFM2-350M NVFP4A16 hits 1.7M tok/s decode on a single RTX 5090

AlpinDale · x · 2026-09-04

The Localmaxxing leaderboard lists Liquid AI's LFM2-350M quantized to NVFP4A16, reaching roughly 1.7M tok/s decode on a single RTX 5090 (32 GB). AlpinDale quipped that at that speed you "might as well read /dev/urandom." The tweet is just the quip plus a link; detailed latency and VRAM numbers live on the linked leaderboard page.

Original post →

More from Infra

Infra channel →