Running Qwen 2.5 on Mac Studio: Performance Tests and Agent Struggles
Over_Technology_1764 · reddit · 2026-08-16
A user tested Qwen 2.5 models (referencing 27B and 72B, despite a typo in the original text) on a Mac Studio with 64GB RAM. Using the community MLX 4-bit quantization yielded a speed of about 16 tps. When attempting to create a Pokémon game with Pi Agent, the model got stuck in a thinking loop and failed to progress with complex tasks. After unsuccessful optimization attempts with Hermes Agent, the user reverted to an older model and is seeking advice from others with similar setups.
Related event: Running Local Agent Models on Mac: Memory Is the Deciding Factor(2 posts)→
More from Infra
- Docker Model Runner supports CNCF ModelPack open standard — HowDevelop · 2026-08-16
- Koboldcpp local inference tool released v1.119 — Fcking_Chuck · 2026-08-16
- All-Optical General-Purpose CPU: Path to Zettaflop Computing? — jwt0625 · 2026-08-16
- Compiler turns HF checkpoints into RTL/GDS in 7k lines — teortaxesTex · 2026-08-16
- Will the data center boom last for decades? — Aurora_7021 · 2026-08-16
- Move data centers to the moon? Let AI ruin an uninhabited planet — DevToD4 · 2026-08-16