Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options
DrainBramage · reddit · 2026-10-06
A Reddit user asks for help choosing a local LLM + agent stack for a new 128GB M5 Max Mac Studio, meant to run coding, browser automation and multi-step agents on sensitive client data without monopolizing all RAM.
Candidates include Hermes Agent, Qwen3.8-Flash-Next, an MTPLX optimized speed build, and Tailscale for remote access. The main confusion is the inference/server layer: LM Studio vs Ollama vs MLX/llama.cpp vs MTPLX, and whether MTPLX replaces or underlies them, plus whether the Flash-Next build is mature enough for daily business use.
More from coding & agent
- Tern, a Rust-native 'neoterminal', hits feature-complete with persistent agent sessions and 150ms startup — sull · 2026-10-06
- GTA 5 ported to WebAssembly with AI's help — 'the end of PC ports?' — pvncher · 2026-10-06
- How Rippling shipped production AI in 6 months with Deep Agents and LangSmith — LangChain · 2026-10-06
- COLM hosts Self-Improving Agents social with talks on Meta-Harness, EvoSkill and agent fragility — tuvllms · 2026-10-06
- Microsoft AI team shares talk on fine-tuning models for knowledge work like Excel — marlene_zw · 2026-10-06
- How do you change a production AI agent's authority without redeploying it? — BaraSlim · 2026-10-06