AI agents rewrite inference engine, boosting a 27B model from 66 to 580 tok/s on Mac
a300a300 · reddit · 2026-09-28
According to the post, AI agents (mostly Opus 5.5) rewrote a 27B model's inference engine over three days, speeding it up from 66 tok/s to 580 tok/s (9x) on a Mac via MLX. Details at yukon.org/mlxfast. If it holds up, this is a striking case of AI optimizing its own inference infrastructure, though specifics remain to be verified.
More from coding & agent
- Codex builds whole apps but keeps choking on reusing your live Chrome session — ChrisGPT · 2026-09-28
- DeepSeek ships Codex-style agent harness that connects to multiple model providers — ChrisUniverse · 2026-09-28
- Indie dev builds CoIsland, a notch-resident AI monitor with 37 connectors, entirely with Claude Code — DonTizi · 2026-09-28
- Andrew Carr says Astra gave bit-identical, hash-verified outputs on his weekend project — andrew_n_carr · 2026-09-28
- SkillGym fine-tuning lifts Qwen3.5 35B past Claude Sonnet 4.6 on agentic coding benchmarks — dair_ai · 2026-09-28
- Coinbase CEO Brian Armstrong: every backend service needs a /feedback endpoint for agents — pzakin · 2026-09-28