Llama.cpp deep dive: Heterogeneous GPU setup boosts speed by 70% and enables 262k context
fintip · reddit · 2026-08-20
After 3 days of benchmarking on a heterogeneous 40GB VRAM setup (RTX 4090 laptop + RX 7900 XTX eGPU), the author optimized llama.cpp parameters to increase generation speed from 16 to 27 t/s and unlock the full 262k context. Key findings include: enabling speculative decoding (MTP + ngram) significantly boosts generation; the 7900 XTX excels at decoding tasks; TB4 bandwidth overhead is manageable; and a bug in llama.cpp's multi-GPU MTP implementation was identified and reported. The optimized launch command is provided.
More from Infra
- Omarchy and Linux Desktop: Opportunities Fueled by AI — antirez · 2026-08-20
- Microsoft migrates TypeScript repo to Go in major PR merge — wateriscoding · 2026-08-20
- Investor Applies Munger's Three-Basket Rule: Data Centers Are a Clear Yes — RachelVT42 · 2026-08-20
- Benchmark: FP8 Models Run 5x Faster Than GGUF on Low-End Hardware — ROBOTTTTT13 · 2026-08-20
- MiniMax H3 + Qwen Image Edit recreates Sherlock shots, 10s per gen on a 3090 — nikhilprasanth · 2026-08-20
- Ops Horror Stories: Redesigning Storage Layout on a Live Ceph Cluster Invites Chernobyl Jokes — l4rz · 2026-08-20