M1 Ultra 128GB Test: Patch Boosts Local DeepSeek V4 to 16 tok/s

mil_phickelson · reddit · 2026-08-03

A developer ran DeepSeek V4 locally on an M1 Ultra 128GB machine via LM Studio (using Unsloth UD-IQ3XXS quantization). By applying a community-provided engine patch, inference speed jumped from 5-6 tok/s to 15-16 tok/s, alongside noticeable improvements in output quality.

Original post →

More from Infra

Infra channel →