AMD 780M iGPU Runs 35B LLMs: Sub-€1000 Local Inference Setup
MaximusSenior · reddit · 2026-08-09
A Reddit user shares a budget-friendly local LLM inference setup using AMD Ryzen 7 260 (or similar) CPUs with 780M iGPU and 64GB DDR5 RAM, running llama.cpp with Vulkan on Ubuntu.
- Hardware cost: barebone mini PC €300-400, used 64GB DDR5 €500, total well under €1000.
- Kernel params like amdgpu.gttsize=49152 map system RAM as VRAM, yielding 48GB "VRAM".
- Performance (Q8 quant): Qwen 3.6 35B-A3B generates 21 t/s; Gemma 4 31B 2.5 t/s, but with MTP reaches 5.76 t/s.
- Bonus: adding a small discrete GPU (e.g., RTX 5060 8GB) can boost MoE models via partial expert offloading.
More from Infra
- Edge Caching for AI Agents: Return High-Frequency Requests Directly at the Edge — blaizedsouza · 2026-08-10
- AMD Acquires Taalas: Etching Model Weights Directly Into Silicon to End GPU Era? — julsimon · 2026-08-09
- AI Demand to Cause Severe HBM Memory Shortage in Next Two Years — RihardJarc · 2026-08-09
- Running DeepSeek v4 Flash Locally on CPU: RTX 4090 + Tesla P40 Setup — DigiDecode_ · 2026-08-09
- NIMBY AI: Money, Power, and Populism in the Data Center Buildout — ruthstarkman · 2026-08-09
- Cathie Wood: US Battery Storage Hits 52 GW, Growing 70% Yearly — CathieDWood · 2026-08-09