Qwen3.8 27B at 11.7 tok/s on RTX 4070 Ti + Mac Air

zannix · reddit · 2026-08-24

User runs Qwen3.8 27B via llama.cpp RPC across an RTX 4070 Ti and an M5 MacBook Air, achieving 11.65 tok/s with 32k context, Q8 KV cache, and MTP enabled. Seeking configuration advice (tensor split, builds) to push towards 15 tok/s without sacrificing accuracy or context length.

Original post →

More from Infra

Infra channel →