DwarfStar runs DeepSeek V4 Flash on M3 Ultra at 37 t/s
mishig25 · x · 2026-08-03
DwarfStar, an inference engine developed by antirez, runs DeepSeek V4 Flash 0731 (mxfp4 quantization) on an M3 Ultra with 512GB of memory, achieving an inference speed of 37 tokens/s.
More from Infra
- TensorSharp Benchmark: Speculative Decoding Doubles DeepSeek Speed — fuzhongkai · 2026-08-03
- Developer Builds MCP Server Covering 4.8M Podcasts and 131M Episodes with SQL Query — Harj0t1singh · 2026-08-03
- Token Efficiency is the Real Bottleneck, Not LLM Routing — AdditionalWeb107 · 2026-08-03
- As Chip TDP Hits 500W, Cooling Challenges for Next-Gen Rubin Surface — jwt0625 · 2026-08-03
- Challenges of scheduling local AI compute across 5090, Mac Studio, and eGPUs — ShittyMillennial · 2026-08-03
- Why SQLite's Architecture is Poised to Dominate the Agentic Era — glcst · 2026-08-03