12,000 tok/s inference achieved on a $250 FPGA board
ycombinator · x · 2026-08-17
Lamb Labs demonstrated a proof-of-concept achieving 12,000 tokens per second inference on a $250 AMD KV260 FPGA board. The project, named "little lamb," runs a small model entirely within FPGA fabric with resident weights in on-chip memory, eliminating DRAM from the token loop. The authors note that while the coherence is poor and it is not intended for good chat results, the extreme speed showcases the potential for specialized hardware acceleration.
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17