Forked Strata hits 7,357 tok/s prefill on IBM AC922 running an 8B Qwen model

okoyl3 · reddit · 2026-10-05

A developer forked the Strata inference engine and, iterating heavily with Claude Code, tuned it for an IBM AC922 (dual POWER9 + 4x V100 SXM2, NVLink unified memory) running an 8B Qwen Q4KXL model. Prefill jumped from llama.cpp's 130 tok/s to a peak 7,357 tok/s (still 7,090 at a 252K-token prompt), decode 113 tok/s via MTP speculative decoding, with all 72 GiB of expert weights page-locked in RAM and GPUs pulling 70 GB/s over NVLink. Open-sourced as Strata-AC922.

Original post →

More from Infra

Infra channel →