Open-source MLX.fast project claims up to 3x faster small-model inference on Apple Silicon
VagabondTruffle · reddit · 2026-09-27
A developer released ishizuki, an open-source MLX-based inference optimization for Apple Silicon, claiming up to 3x speedups for small models like Qwen3.8+Flash-Next on Mac chips. The author previously topped the MLX.fast leaderboard and says the implementation remains the fastest on chips below M5, inviting community contributions.
More from Infra
- Polymarket puts 29% odds on any US state enacting a data center moratorium this year — Polymarket · 2026-09-27
- Big 4 AI capex hits 2.4% of GDP, double the telecom bubble peak — but still below railroads — rickasaurus · 2026-09-27
- Goldman: token demand to grow 18x by Sept 2026, but frontier demand lags at 8-9x — rohanpaul_ai · 2026-09-27
- Banks price orbital compute at $165-180B per GW; new model says $68B by 2028 — JOBhakdi · 2026-09-27
- Modal open-sources GPU Glossary, a human-friendly dictionary for GPU programming terms — charles_irl · 2026-09-27
- MiMo V2.6 Flash ported to Windows sglang on a 6-GPU 5090/3090 cluster — comperr · 2026-09-27