Indie run: 23k tps decode on a single RTX 5090, matching Llama 3.2 1B at ~90% less cost per token

jon_durbin · x · 2026-10-03

Developer jondurbin shares preliminary results from his Kappa training run (Gated DeltaNet-2, 576B checkpoint, 9T tokens): it beats Llama 3.2 1B on ARC-C, OBQA, TQA and nearly matches it on ARC-E, SciQ, at roughly 90% lower cost per token. His Parallax vLLM fork hits 23k tps decode at high batch on a single RTX 5090 (28k with fp16 GDN2) and 70k tps prefill; llama.cpp fork, mobile inference and training code will be open sourced soon. Known issues include an underflow in a GDN2 per-channel decay gate.

Related event: Open Pretraining Run Matches Llama 3.2 1B at One-Tenth the Cost(2 posts)→

Original post →

More from Infra

Infra channel →