Testing HY3 on a 128GB Mac
amerkthetrippyone · reddit · 2026-07-11
This post shares hands-on experience running Tencent's HY3 model on a MacBook M5 Max 128GB, focusing on quantized files, llama.cpp support, and performance.
Main Points
- The model introduced is the open-weights version of 295B-A21B MoE, which the author considers cutting-edge among open models.
- A 107GB UD128 quantization was chosen because the repository provides public perplexity data, allowing at least a rough estimate of quantization loss.
- To run it, the author switched to a llama.cpp PR supporting HY3 and increased the Mac's GPU memory limit.
- Due to a mismatch between the architecture name in the GGUF metadata and the registered name in the PR, the author had to manually edit the model file fields to launch it successfully.
Related event: Tencent HY3 Model Successfully Deployed on 128GB Mac(2 posts)→
More from Infra
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22
- SkyPilot comes out of stealth with a pitch to unify fragmented AI compute — skypilot_org · 2026-07-22
- SkyPilot emerges from stealth with over $20M to tackle fragmented AI compute — skypilot_org · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- AI agent accountability layer adds terminal verification with explicit finality and no signup — Special_Librarian145 · 2026-07-22