llama.cpp VRAM Spikes with Long Contexts on Intel Arc B580
WizardlyBump17 · reddit · 2026-08-14
A developer running llama.cpp locally via SYCL on an Intel Arc B580 GPU encountered a VRAM bottleneck. While the model consumes 10.7GB of VRAM upon loading, sending a prompt immediately spikes usage to 11.6GB. As the conversation context grows, the VRAM usage eventually fills up the entire 12GB and can even cause a crash.
The author noted that switching to the Vulkan backend keeps VRAM usage flat, but at the cost of significantly slower inference speeds. They shared their Docker configuration running a Qwen model and asked the community if this VRAM scaling issue is a common trait across Nvidia and AMD cards, seeking advice on optimizing long-context handling.
More from Infra
- Inside Cerebras WSE: 900k Cores and an Alien Kernel Programming Model — i_dg23 · 2026-08-14
- Open Source Models Are Getting Bigger, But Consumer GPU is the Bottleneck — flowersslop · 2026-08-14
- The 'Compute Dollar' Will Replace the Petrodollar and Define the Next 50 Years — NinaDSchick · 2026-08-14
- Running Qwen 2.4T Locally: 5 GPUs Still Can't Make It Viable — klicker0 · 2026-08-14
- Orange Pi Launches AI Station: 176 TOPS NPU and up to 96GB RAM — MundanePercentage674 · 2026-08-14
- xAI Reveals Memphis Supercomputer Site Has Paid $30M in Taxes — elonmusk · 2026-08-14