Running Two Models Across Strix Halo + r9700 Hits OOM: Full Config Shared

El_90 · reddit · 2026-09-14

A Redditor runs Laguna-S 2.1 Q6 on Strix Halo 128GB (Vulkan1, 107G/124G GTT used) and Qwen3.8-27B Q6 on an oculink-attached r9700 (Vulkan0). Each works alone, but loading both triggers a kernel page allocation OOM in llama-server. They've tried lowering context, tuning pagepoolsize, DMA32, and kernel params (amdiommu=off, ttm.pageslimit), and shared full llama.cpp launch configs, asking what non-memory resource llama.cpp might be exhausting.

Original post →

More from Infra

Infra channel →