Vulkan runs 20°C cooler than CUDA on laptops in llama.cpp, with a catch

Hot-Employ-3399 · reddit · 2026-08-28

A thermal-conscious laptop user reports that in llama.cpp on a 27B Qwen model, the Vulkan backend runs 20°C cooler than CUDA (65-75°C vs 85-95°C on long generations) while also gaining 2-3 tokens/s. Downsides: model loading takes forever, and with partial layer offload (e.g., a 35B MoE) throughput collapses below 10 tok/s. Limiting GPU clocks or lowering thread count changed nothing. He's asking for better settings or whether vLLM/exllama alternatives are worth testing.

Original post →

More from Infra

Infra channel →