Running Two Models Across Strix Halo + r9700 Hits OOM: Full Config Shared
El_90 · reddit · 2026-09-14
A Redditor runs Laguna-S 2.1 Q6 on Strix Halo 128GB (Vulkan1, 107G/124G GTT used) and Qwen3.8-27B Q6 on an oculink-attached r9700 (Vulkan0). Each works alone, but loading both triggers a kernel page allocation OOM in llama-server. They've tried lowering context, tuning pagepoolsize, DMA32, and kernel params (amdiommu=off, ttm.pageslimit), and shared full llama.cpp launch configs, asking what non-memory resource llama.cpp might be exhausting.
More from Infra
- Attestable Builds a Verification Layer to Check If Compute Follows AI Rules — ml_hardware · 2026-09-15
- Forge launches with Arcee, Microsoft, Vercel to push open weight models to the frontier — inkko44 · 2026-09-15
- Polymarket prices 13% odds of AI bubble bursting by end of 2026, with strict resolution rules — Polymarket · 2026-09-14
- François Fleuret: MiniMax runs well on a $500 16GB 4060Ti when you prompt per the docs — francoisfleuret · 2026-09-14
- The power gap is an intelligence gap: why Europe's grid threatens its AI future — NinaDSchick · 2026-09-14
- Cloud Run creator Steren joins Vercel to lead Fluid compute products — RealGeneKim · 2026-09-14