Optimize llama.cpp Dual-GPU Tensor Split for Larger Context
ea_man · reddit · 2026-08-31
The author shares a method to fine-tune VRAM allocation in llama.cpp across two GPUs using --override-tensor. Unlike rough ratios, this approach moves specific tensors to maximize space, increasing context length for a Qwen model from 110k to 139k tokens. A script tool is provided to automate the probing process on Linux (supporting ROCm/Vulkan), along with a workflow involving KV cache saving to validate the results.
More from Infra
- drskill released: A 'brew doctor' for checking Agent context and MCP configs — dbreunig · 2026-08-31
- SK Hynix expects 60-100% surge in AI memory demand by 2027 — Beth_Kindig · 2026-08-31
- Carnot & Abacus: Enterprise Agent Query Systems with Cost Optimization — lateinteraction · 2026-08-31
- Rust-based visloc-rs library outperforms COLMAP in SfM accuracy — rsasaki0109 · 2026-08-31
- Google data center deal in Georgia could eliminate property taxes for residents — robleclerc · 2026-08-31
- Google, Amazon, and others fund utility to cut Ohio residents' electricity bills — SydSteyerhart · 2026-08-31