Custom CPU engine for Qwen3.5 0.8B beats llama.cpp with 2.9x faster prefill

Danmoreng · reddit · 2026-09-08

Reddit user Danmoreng vibecoded a small C++ inference engine and a custom 4-bit quant format (H128/Q4-G32-DOT4, with activation-based calibration and blockwise error compensation) to run Qwen3.5 0.8B on CPU — his use case is local dictation cleanup while the GPU is busy.

Original post →

More from Infra

Infra channel →