DwarfStar runs DeepSeek V4 Flash on M3 Ultra at 37 t/s

mishig25 · x · 2026-08-03

DwarfStar, an inference engine developed by antirez, runs DeepSeek V4 Flash 0731 (mxfp4 quantization) on an M3 Ultra with 512GB of memory, achieving an inference speed of 37 tokens/s.

Original post →

More from Infra

Infra channel →