Forked Strata hits 7,357 tok/s prefill on IBM AC922 running an 8B Qwen model
okoyl3 · reddit · 2026-10-05
A developer forked the Strata inference engine and, iterating heavily with Claude Code, tuned it for an IBM AC922 (dual POWER9 + 4x V100 SXM2, NVLink unified memory) running an 8B Qwen Q4KXL model. Prefill jumped from llama.cpp's 130 tok/s to a peak 7,357 tok/s (still 7,090 at a 252K-token prompt), decode 113 tok/s via MTP speculative decoding, with all 72 GiB of expert weights page-locked in RAM and GPUs pulling 70 GB/s over NVLink. Open-sourced as Strata-AC922.
More from Infra
- Local inference in practice: mining repos, AI chat logs and media libraries with a tiny model — natesiggard · 2026-10-05
- Stripping antirez's ds4 to 45k lines makes Qwen3.8 Flash Next ~10% faster, bit-exact — Chida82 · 2026-10-05
- Hugging Face goes down on final day of paper submissions — silver__tsuki · 2026-10-05
- Running Qwen3.5 9B/27B INT4 on cheap ex-mining FPGA boards — I_am_purrfect · 2026-10-05
- DLSS5 hands-on: neural rendering delivers '5 years of graphics progress in one toggle' — ctnzr · 2026-10-05
- 800G/1.6T DSP supply tight across vendors; MaxLinear 1.6T chip still sampling — iamfabian · 2026-10-04