75-test coding benchmark pits RTX 4090 vs Strix Halo vs M5 Max for local LLMs
julianharris · x · 2026-10-11
The author is benchmarking three local AI machines — RTX 4090/64GB, Minisforum Strix Halo 128GB, and MacBook Pro M5 Max 128GB — with Qwen models and Flash Next.
- Method: build the same 'MVP Miro clone' app, 11 stories, 75 holdout tests, 5 runs each for statistical rigor
- Some combinations take over 24h for a single run
- Toolchain spans llama.cpp, mlx-serve, Unsloth, quantization schemes, and multiple coding agent frontends
- Results not yet published, but long soak tests have surfaced many bugs
More from coding & agent
- Prime Agent orchestrates 2,000+ agent swarm to rewrite itself in Rust, 13x faster startup — GregCook2011 · 2026-10-11
- Dario Amodei predicted software engineering automatable in 12 months — tech leaders say it's coming true — matthew_d_green · 2026-10-11
- Regents Labs runs a daily AI paper digest with ChatGPT Pro, spotlighting Prime Intellect's agent stack — seanwbren · 2026-10-11
- Building an Agent Feedback Loop to Auto-Tune Game Difficulty via Headless Simulation — zimmer550king · 2026-10-11
- Python vs Go vs Rust for production agentic systems: a six-dimension practical evaluation — mostly_deterministic · 2026-10-11
- Agent Use Cases: a free site scanning the web daily for 1,000+ real AI agent use cases — Roger_M_Taylor · 2026-10-11