Pure MLX Engine Hits 65 tok/s, Cuts Model Size in Half
EyalToledano · x · 2026-08-28
The author is building a pure MLX inference engine to bypass dependency constraints and maximize machine performance. Key wins so far include:
- Speed: Reached 65 tok/s with zero accuracy loss, aiming for 100+.
- Memory: Expert paging reduces model size from 39GB to 20GB.
Benchmarks show Qwen3.8-Flash-Next-REAP variants maintain high HumanEval scores while significantly reducing resident memory.
More from Infra
- Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090 — _reachsumit · 2026-08-28
- Intel XE3P projected specs: 1.3 PFLOPS FP8, 1.5TB/s bandwidth, 2027 launch — QuixiAI · 2026-08-28
- Rumor: Anthropic interested in developing its own training chip — zephyr_z9 · 2026-08-28
- India Commits $13.4B for 'Semicon 2.0' Chip Design and Manufacturing — SumitGup · 2026-08-28
- Meta, Google, NTT to discuss AI data center optical architectures — jwt0625 · 2026-08-28
- Coherent-lite Catches Up to IMDD in Energy Efficiency for Pluggables — jwt0625 · 2026-08-28