Running Qwen3.5 9B/27B INT4 on cheap ex-mining FPGA boards

I_am_purrfect · reddit · 2026-10-05

Background

The author wanted FPGA LLM inference for a while; Qwen3.5 finally made the 9B-27B scale compelling. Cheap hardware: SQRL FK33 ($280, 8GB HBM2, 400GB/s, eBay ex-mining card), later the dual-VU35P Jungle Cat ($375). The RTL was largely implemented with Claude Opus 4.8/5.5 and Kimi K3.

Measured results

ASIC estimate

An Opus 5.5 estimate puts this RTL on TSMC 2023 N3 with 6 stacks of HBM3 (4.9TB/s) at 2GHz at roughly 294 tok/s short-context, 106 tok/s at 16k, 10 tok/s at 262k for the 27B, at 125-340W. Two BC-250s bought for $60/$75 proved an incredible value.

MIT-licensed repo: llm.vhdl

Original post →

More from Infra

Infra channel →