DeepSeek V4 Flash on M2 Ultra: lossless repack to 141GiB, 25.8 t/s beats M3 Ultra
Agusx1211 · reddit · 2026-08-23
A developer built a custom llama.cpp fork for the M2 Ultra (60 cores, 192GB) that repacks DeepSeek V4 Flash losslessly to 141 GiB — smaller than the public Q4 GGUF (freeing room for context), with byte-identical output and no KV cache quantization.
Key results:
- 25.8 t/s output (42 t/s peak), faster than the 16 t/s seen on M3 Ultra
- SSD KV cache and dynamic lanes, 1M total context across 8 lanes
- Prompt processing is low (250 t/s at 8k-32k) but the SSD cache compensates heavily; the author says compute is still left on the table
Open-sourced at llama-cpp-ds4f-m2-ultra.
More from Infra
- Cursor Launches S3-Based Git Storage System for Scale — bibryam · 2026-08-23
- PromoteOps: MCP Server for Automating AWS CloudFormation Promotion — iamthanoss · 2026-08-23
- NAS boot failure复盘: Store metadata on separate drives — TheZachMueller · 2026-08-23
- Nvidia Rubin NVL72 Rack Could Cost ~$8M — zephyr_z9 · 2026-08-23
- Vertex AI Users Complain: No Native Hard Spending Cap, Must DIY via Pub/Sub — MeetingWeird9418 · 2026-08-23
- Open source MCP server implements HTTP 402 micropayment gateway — EstablishmentTough18 · 2026-08-23