Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options

DrainBramage · reddit · 2026-10-06

A Reddit user asks for help choosing a local LLM + agent stack for a new 128GB M5 Max Mac Studio, meant to run coding, browser automation and multi-step agents on sensitive client data without monopolizing all RAM.

Candidates include Hermes Agent, Qwen3.8-Flash-Next, an MTPLX optimized speed build, and Tailscale for remote access. The main confusion is the inference/server layer: LM Studio vs Ollama vs MLX/llama.cpp vs MTPLX, and whether MTPLX replaces or underlies them, plus whether the Flash-Next build is mature enough for daily business use.

Original post →

More from coding & agent

coding & agent channel →