Vulkan runs 20°C cooler than CUDA on laptops in llama.cpp, with a catch
Hot-Employ-3399 · reddit · 2026-08-28
A thermal-conscious laptop user reports that in llama.cpp on a 27B Qwen model, the Vulkan backend runs 20°C cooler than CUDA (65-75°C vs 85-95°C on long generations) while also gaining 2-3 tokens/s. Downsides: model loading takes forever, and with partial layer offload (e.g., a 35B MoE) throughput collapses below 10 tok/s. Limiting GPU clocks or lowering thread count changed nothing. He's asking for better settings or whether vLLM/exllama alternatives are worth testing.
More from Infra
- Anthropic discussed $7B MatX acquisition; startup now raising at $4B valuation — DZhang50 · 2026-08-28
- Three 32GB AMD R9700s for the price of one RTX 5090: a local LLM user's dilemma — MrHall · 2026-08-28
- DeepSeek V4 Flash pricing puzzle: How does it match GPT OSS 20B? — gajesh · 2026-08-28
- Qwen3.8-Flash reportedly costs 1/9th to train vs Qwen3.7-Plus — VraserX · 2026-08-28
- Sovereign AI market hits $1.5T as companies flee US cloud providers — mikeflache · 2026-08-28
- ComfyUI multi-GPU setups: text+VAE on one card, diffusion on the other — hurdurdur7 · 2026-08-28