Nex-N2.5-mini-MLX-4bit hits 133.6 tok/s on Apple M5 Max

DerTomsn · reddit · 2026-09-13

Benchmarks of the newly released Nex N2.5 Mini (MLX 4bit quant) on Apple M5 Max show 133.6 tok/s generation with recommended settings (temp 0.7, topp 0.95, topk 40, reasoningeffort high), plus fast prompt processing, reasonable memory footprint, and solid quality. Full results are on llm-bench.io, quant on Hugging Face. Author plans to use it for agents and coding.

Original post →

More from Infra

Infra channel →