Running Qwen3.5 9B/27B INT4 on cheap ex-mining FPGA boards
I_am_purrfect · reddit · 2026-10-05
Background
The author wanted FPGA LLM inference for a while; Qwen3.5 finally made the 9B-27B scale compelling. Cheap hardware: SQRL FK33 ($280, 8GB HBM2, 400GB/s, eBay ex-mining card), later the dual-VU35P Jungle Cat ($375). The RTL was largely implemented with Claude Opus 4.8/5.5 and Kimi K3.
Measured results
- Qwen3.5-9B INT4, 2x FK33 @75MHz pipeline split: 6 tok/s prefill (256-token prompt), 3.2 tok/s generation early, 2.4 tok/s at 2-3k context; outputs verified layer-by-layer against llama.cpp
- Qwen3.8-27B INT4 (modelled from 9B per-op profile, not yet run): 1.1 tok/s generation with current design on two dies; resized RTL at 200MHz 8 tok/s; 4x VU35P tensor-parallel at 200MHz estimated 25 tok/s short-context, 10 tok/s at 16k, 1 tok/s at 262k. Two dies cap out around 45k context (KV cache won't fit beside 14.5GB of weights); full 262k needs four dies.
ASIC estimate
An Opus 5.5 estimate puts this RTL on TSMC 2023 N3 with 6 stacks of HBM3 (4.9TB/s) at 2GHz at roughly 294 tok/s short-context, 106 tok/s at 16k, 10 tok/s at 262k for the 27B, at 125-340W. Two BC-250s bought for $60/$75 proved an incredible value.
MIT-licensed repo: llm.vhdl
More from Infra
- Local inference in practice: mining repos, AI chat logs and media libraries with a tiny model — natesiggard · 2026-10-05
- Stripping antirez's ds4 to 45k lines makes Qwen3.8 Flash Next ~10% faster, bit-exact — Chida82 · 2026-10-05
- Hugging Face goes down on final day of paper submissions — silver__tsuki · 2026-10-05
- DLSS5 hands-on: neural rendering delivers '5 years of graphics progress in one toggle' — ctnzr · 2026-10-05
- Forked Strata hits 7,357 tok/s prefill on IBM AC922 running an 8B Qwen model — okoyl3 · 2026-10-05
- 800G/1.6T DSP supply tight across vendors; MaxLinear 1.6T chip still sampling — iamfabian · 2026-10-04