Benchmarks show Gufo's 70 tok/s claim only holds on repeat-a-word prompts; halogen is faster in practice

brainchillzZ · reddit · 2026-10-02

A hands-on benchmarking post dissects open-source inference engine Gufo's headline "70.56 tok/s on Qwen 27B Q4" on a Strix Halo device. The author reproduced the number but found it relies on a "write red 1000 times" prompt where the speculative-decoding draft model is always right; on nine normal prompts throughput drops to a median 39.4 tok/s (22–52 range). The "123 tok/s aggregated" 8-user figure also omits queue time—actual delivered tokens are 82/52.

Head-to-head with halogen 0.13.8 on identical hardware: halogen is 13% faster single-user (43.9 vs 38.2) and 18% faster with 4 users (76.6 vs 63.1), while Gufo wins on long-prompt processing (1495 vs 1288 tok/s) and repetitive prompts. To its credit, Gufo's docs do separate mixed vs repetitive numbers—the spin is only in the README—and 39 tok/s without quality comparison is still strong for a Q4 27B on an APU. Non-performance upsides: plain GGUF support, MIT license, 3s model load, per-request draft/cache logging, 8-session batching, plus ASR/TTS/image features.

Original post →

More from Infra

Infra channel →