How to Allocate VRAM on Strix Halo
mindwip · reddit · 2026-07-11
The post asks how to properly allocate VRAM when running local models on machines like the Strix Halo: should it be 1GB, 512MB, or scaled up to 96GB? The poster notes some claim leaving most memory for the system runs just as fast, though they recall "more VRAM equals faster."
Core questions include:
- What is the optimal VRAM vs. system memory allocation for Strix Halo on Windows / Linux?
- To run larger models, is it better to lean towards "minimal VRAM + more system memory" or "max VRAM allocation"?
- Are there performance discrepancies on AMD platforms between GPU and CPU cores, or VRAM and system memory?
More from Infra
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21
- Former AWS operator says Bedrock margins can beat SageMaker as agentic AI lifts CPU demand — RihardJarc · 2026-07-21