Quantized MacBook Nearly Matches DGX Spark
anvarazizov · reddit · 2026-07-18
The author used Terminal-Bench 2.1 and Terminus-2 to compare two local systems running DeepSeek-V4-Flash:
- 1 M5 Max MacBook: Highly quantized 80.8 GiB GGUF at roughly 2.45 bit/weight
- 2× DGX Spark: Native FP8/FP4 checkpoint with DSpark speculative decoding
Results show:
- MacBook: 47/87 = 54.0%
- 2× DGX Spark: 45/86 = 52.3%
- Out of the 86 tasks scorable on both sides, they reached identical conclusions on 66 of them
- There were 20 disputed tasks: 11 Mac-only and 9 Spark-only
The author emphasizes that this does not prove 2-bit quantization is a "free lunch" because the variables were not strictly controlled: hardware, KV format, context limits, runtime, and Spark's speculative decoding all differed. More accurately, this is an end-to-end comparison of two "complete systems" rather than a pure quantization ablation.
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11