Nex-N2.5-mini-MLX-4bit hits 133.6 tok/s on Apple M5 Max
DerTomsn · reddit · 2026-09-13
Benchmarks of the newly released Nex N2.5 Mini (MLX 4bit quant) on Apple M5 Max show 133.6 tok/s generation with recommended settings (temp 0.7, topp 0.95, topk 40, reasoningeffort high), plus fast prompt processing, reasonable memory footprint, and solid quality. Full results are on llm-bench.io, quant on Hugging Face. Author plans to use it for agents and coding.
More from Infra
- MHA, MQA, GQA and MLA explained by what happens to the K/V cache during decoding — techNmak · 2026-09-13
- Early OpenAI employee says 'winning' AGI is outdated — 99% of future compute will run locally — GregCook2011 · 2026-09-13
- Benchmarks show ComfyUI in Docker (CUDA 12.4) runs at 0% penalty if you fix the /dev/shm OOM crash — fluxdraw · 2026-09-13
- Qwen3.8-27B EXL3 one-click kit brings quality local LLM to 16-32GB consumer GPUs — udmrzn · 2026-09-13
- Positron shipped an AI chip in 15 months; Ainek rumored at 13 — DavidBennett__ · 2026-09-13
- nic_carter rebuilds Meta datacenter scale graphic from official schematics — AccBalanced · 2026-09-13