12,000 tok/s inference achieved on a $250 FPGA board

ycombinator · x · 2026-08-17

Lamb Labs demonstrated a proof-of-concept achieving 12,000 tokens per second inference on a $250 AMD KV260 FPGA board. The project, named "little lamb," runs a small model entirely within FPGA fabric with resident weights in on-chip memory, eliminating DRAM from the token loop. The authors note that while the coherence is poor and it is not intended for good chat results, the extreme speed showcases the potential for specialized hardware acceleration.

Original post →

More from Infra

Infra channel →