M1 Max Benchmarks: 72 tok/s Aggregate Throughput at 128k Context
EyalToledano · x · 2026-09-02
Benchmarks for the M1 Max (64GB RAM) show better performance than the M1 Pro due to higher memory bandwidth (400 GB/s). With Q4 quantization, 128k context, and 100% experts in RAM, it achieves 46 tok/s burst and 72 tok/s aggregate throughput with 6 agents. However, memory usage approaches the limit at 256k context.
More from Infra
- Intel exec: AI era security requires silicon-level design, not afterthoughts — BenBajarin · 2026-09-02
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- Exllamav3 benchmarks: 700tk/s on 8x3090 setup — Leflakk · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02
- Asus AI PC Price Jumps 50%, Speculating on Upcoming DGX Spark Hike — mountainyoo · 2026-09-02
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02