M1 Ultra 128GB Test: Patch Boosts Local DeepSeek V4 to 16 tok/s
mil_phickelson · reddit · 2026-08-03
A developer ran DeepSeek V4 locally on an M1 Ultra 128GB machine via LM Studio (using Unsloth UD-IQ3XXS quantization). By applying a community-provided engine patch, inference speed jumped from 5-6 tok/s to 15-16 tok/s, alongside noticeable improvements in output quality.
More from Infra
- Chrome Canary Introduces Native Embedding API for On-Device AI — gaganghotra_ · 2026-08-03
- Decentralized Compute Network Surges to 6,000 GPUs in Two Weeks — 0xSammy · 2026-08-03
- China's DFSX System Claims 2x Memory Bandwidth of NVIDIA's GB200 Using Vertical Compute Memory Towers — MundanePercentage674 · 2026-08-03
- llama.cpp Launches Official Mac App and One-Command Server Setup — rm-rf-rm · 2026-08-03
- CoreAutoAI Event: Building the World's Most Automated AI Lab & System Optimizations — marksaroufim · 2026-08-03
- Meta's Secret High-Performance GPU Kernel Library MSLK Documented by AI Agents — giffmana · 2026-08-03