Quantized MacBook Nearly Matches DGX Spark
anvarazizov · reddit · 2026-07-18
The author used Terminal-Bench 2.1 and Terminus-2 to compare two local systems running DeepSeek-V4-Flash:
- 1 M5 Max MacBook: Highly quantized 80.8 GiB GGUF at roughly 2.45 bit/weight
- 2× DGX Spark: Native FP8/FP4 checkpoint with DSpark speculative decoding
Results show:
- MacBook: 47/87 = 54.0%
- 2× DGX Spark: 45/86 = 52.3%
- Out of the 86 tasks scorable on both sides, they reached identical conclusions on 66 of them
- There were 20 disputed tasks: 11 Mac-only and 9 Spark-only
The author emphasizes that this does not prove 2-bit quantization is a "free lunch" because the variables were not strictly controlled: hardware, KV format, context limits, runtime, and Spark's speculative decoding all differed. More accurately, this is an end-to-end comparison of two "complete systems" rather than a pure quantization ablation.
More from Infra
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- HilbertRaum open-sources a fully local AI chat and document analysis app for private use — Vladowski · 2026-07-22
- Hybrid and local inference are emerging as a response to AI energy and token costs — dmitry140 · 2026-07-22
- NVIDIA details Vera CPU with 2x performance claims and a 22,000-core rack — ryanshrout · 2026-07-22
- NVIDIA says Vera Rubin NVL72 delivers 10x more tokens per megawatt than Blackwell — nvidia · 2026-07-22