Local LLM on Mac: M2 Ultra 192GB Long-Context Inference Benchmarks
Badger-Purple · reddit · 2026-08-02
A developer tested the local inference performance of a large model (Deepseek-V4-Flash-0731 Dwarfstar) on an M2 Ultra machine with 192GB of unified memory.
Performance Data:
- Prefill phase: Strong performance (see attached chart)
- Decode phase: Starts at 28 t/s; drops to 25 t/s at 45k context length; maintains 17 t/s even at the 192k limit.
More from Infra
- 30B Video MoE Quantizations Tested: Most Users Should Wait — EntireBig7258 · 2026-08-03
- AI Data Centers Consume Up to 1.5 Billion Gallons of Water Yearly — AndyMasley · 2026-08-03
- Prepping for Local LLM Inference: Enthusiast Builds 30TB SSD & 256GB RAM Rig — reto-wyss · 2026-08-03
- CPO Packaging Tech Unlikely to See High-Volume Shipments Before 2028 — BenBajarin · 2026-08-03
- DeepSeek V4 Flash Prefill Speed Boost: Downgrade to CUDA 13.1 — fragment_me · 2026-08-03
- Cornell Releases Roadmap for Parallel Programming and HPC Concepts — thehiphopswami · 2026-08-03