Qwen3.8 27B at 11.7 tok/s on RTX 4070 Ti + Mac Air
zannix · reddit · 2026-08-24
User runs Qwen3.8 27B via llama.cpp RPC across an RTX 4070 Ti and an M5 MacBook Air, achieving 11.65 tok/s with 32k context, Q8 KV cache, and MTP enabled. Seeking configuration advice (tensor split, builds) to push towards 15 tok/s without sacrificing accuracy or context length.
More from Infra
- Hearing language analysis: data center opposition shifts from energy to health and noise — WillRinehart · 2026-08-24
- We gave agents real email addresses and broke deliverability, threading, and privacy — saltexx · 2026-08-24
- GPU Price Hike Favors Early Adopters of Blackwell and Rubin Before 2027 — GavinSBaker · 2026-08-24
- Nvidia NVLink Fusion connects custom XPUs to its AI infrastructure for faster time-to-market — nordicinst · 2026-08-24
- Xiaomi launches new Xring chip, manufactured on TSMC 3nm — pstAsiatech · 2026-08-24
- Napkin math: even a full NY data center ban slows AI by less than a day — random_walker · 2026-08-24