Open-source MLX.fast project claims up to 3x faster small-model inference on Apple Silicon

VagabondTruffle · reddit · 2026-09-27

A developer released ishizuki, an open-source MLX-based inference optimization for Apple Silicon, claiming up to 3x speedups for small models like Qwen3.8+Flash-Next on Mac chips. The author previously topped the MLX.fast leaderboard and says the implementation remains the fastest on chips below M5, inviting community contributions.

Original post →

More from Infra

Infra channel →