Trade a $16K RTX Pro 6000 rig for an M5 Max 128GB MacBook for local inference?
aghanims-scepter · reddit · 2026-08-20
A Reddit discussion: the author runs a 9950x3D + 256GB RAM + RTX Pro 6000 local AI server, but after full-time AI coding at work only uses it for non-coding local inference (private health/finance data). They're considering selling the now-$16K-MSRP card for an M5 Max 128GB MacBook Pro that doubles as laptop and inference box, eliminating always-on server and remote-access headaches.
Key questions: how does M5 Max handle dense models like Qwen3.6/3.8 and Gemma4 at higher context, do 14" models throttle during long inference runs, and how to safely sell an RTX 6000.
More from Infra
- Monad Agent Hub launches with no-code platforms for instant agent creation — bgmshana · 2026-08-20
- AWS Leverages AI Infrastructure Demand to Extend Cloud Dominance — DavidLinthicum · 2026-08-20
- Reverse-Engineering RK3588 NPU: Open Compiler Runs GPT-2 at 36 tok/s — one_does_not_just · 2026-08-20
- llama.cpp dflash2: Qwen 3.8 27B Inference Speed Up to 3x — Top-Eye-8104 · 2026-08-20
- NVIDIA announces FLARE Day event for September — AllThingsApx · 2026-08-20
- Dev billing guide: Pay-per-token may undercut subscriptions — heypearlai · 2026-08-20