Dual Intel Arc Pro B60s Eat Host RAM via GTT — a One-Line Level Zero Workaround
fobsthedev · reddit · 2026-09-21
A Redditor debugged a rare issue with 2× Intel Arc Pro B60 (24GB each) + xe driver + ComfyUI/DisTorch2: GPU allocations mirrored almost 1:1 into host RAM via GTT mappings (8GiB VRAM alloc dropped MemFree by 7559MiB), causing OOM kills that looked like ComfyUI memory shortages. The fix: set NEOReadDebugKeys=1 and ForceZeDeviceCanAccessPerReturnValue=0 so Level Zero's zeDeviceCanAccessPeer() returns false. Host RAM usage fell from multiple GiB to a few hundred MiB while cross-GPU transfers still worked. Debugging used /proc/self/fdinfo and DRM stats; the author seeks kernel 7.1+ users to confirm a kernel-side fix.
More from Infra
- 99.7% cache hits: engineered DeepSeek Harness with self-hosted GLM-5.3 — burny_tech · 2026-09-21
- Solo dev open-sources 4 systems projects, asks engineers to roast them — Accomplished_Row1433 · 2026-09-21
- Andrew Chen: strong LLMs are far from running on phones, on-device AI faces bandwidth, heat and model-size hurdles — andrewchen · 2026-09-21
- Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month — BLUECOW009 · 2026-09-21
- Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP — bring_back_the_v10s · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21