Benchmarks show Gufo's 70 tok/s claim only holds on repeat-a-word prompts; halogen is faster in practice
brainchillzZ · reddit · 2026-10-02
A hands-on benchmarking post dissects open-source inference engine Gufo's headline "70.56 tok/s on Qwen 27B Q4" on a Strix Halo device. The author reproduced the number but found it relies on a "write red 1000 times" prompt where the speculative-decoding draft model is always right; on nine normal prompts throughput drops to a median 39.4 tok/s (22–52 range). The "123 tok/s aggregated" 8-user figure also omits queue time—actual delivered tokens are 82/52.
Head-to-head with halogen 0.13.8 on identical hardware: halogen is 13% faster single-user (43.9 vs 38.2) and 18% faster with 4 users (76.6 vs 63.1), while Gufo wins on long-prompt processing (1495 vs 1288 tok/s) and repetitive prompts. To its credit, Gufo's docs do separate mixed vs repetitive numbers—the spin is only in the README—and 39 tok/s without quality comparison is still strong for a Q4 27B on an APU. Non-performance upsides: plain GGUF support, MIT license, 3s model load, per-request draft/cache logging, 8-session batching, plus ASR/TTS/image features.
More from Infra
- a16z: every $100 into AI buildout sends $50 to chips, $20 to power; 100+ charts — demian_ai · 2026-10-02
- Workload-aware inference: why batch LLM pipelines should plan queries like databases do — sh_reya · 2026-10-02
- Where does GPU spend actually go: training, inference, or idle reserved capacity? — Borges_Engineer · 2026-10-02
- HBM4 controllers eat ~16% of Nvidia's Rubin die, sparking optical interconnect debate — BenBajarin · 2026-10-02
- Modal Clusters goes GA: instant RDMA-connected GPU nodes billed by the second via one decorator — josh_wills · 2026-10-02
- Benchmark Backs Chip Startup Tendrils Compute in Round That Could Top $1B Valuation — Sethwinterroth · 2026-10-02