AI agents rewrite inference engine, boosting a 27B model from 66 to 580 tok/s on Mac

a300a300 · reddit · 2026-09-28

According to the post, AI agents (mostly Opus 5.5) rewrote a 27B model's inference engine over three days, speeding it up from 66 tok/s to 580 tok/s (9x) on a Mac via MLX. Details at yukon.org/mlxfast. If it holds up, this is a striking case of AI optimizing its own inference infrastructure, though specifics remain to be verified.

Original post →

More from coding & agent

coding & agent channel →