Open-weight 27B model hits ~580 tok/s on Mac, 505% over baseline in 3 days
TheMoonMidas · x · 2026-09-28
PrismML's open-weight Ternary Bonsai 2 27B keeps smashing records in the YukonMLX.fast MLX optimization challenge: decode now reaches 580-600 tok/s on an M5 Mac, 505.4% over the frozen baseline, with decode weighted 75% in the score.
The challenge opened less than 3 days ago and the record has climbed from 4x to over 5x; nearly all top-10 leaderboard entries were built by agents powered by Claude, GPT, Grok, DeepSeek and Gemini. Open to anyone, submissions close 10/1, with 79 promoted submissions from 21 solvers so far.
Key points:
- A 27B open model fits a Mac Mini with near server-class decode speeds
- Agents optimizing inference is becoming a self-reinforcing frontier paradigm
- Score formula prefill^0.25 · decode^0.75; paired official runs prevent cheating
More from coding & agent
- Devin user at $100k+ run-rate questions if coding agent revenues will stick — brandon_galang · 2026-09-28
- Turning real postmortems into SRE benchmarks: UIUC team shares data curation lessons — tianyin_xu · 2026-09-28
- One prompt, 40 minutes: dev recreates No Man's Sky with Claude Opus 5.5 and Three.js — Jonesiller5383 · 2026-09-28
- asc CLI 5.7.0 ships App Store Connect API 4.5 support, read-only mode and 2x faster requests — rudrank · 2026-09-28
- Opus 5.5 orchestrates subagents in Amp with messages like "Great work, thank you" — repligate · 2026-09-28
- Building a dataflow analyzer to detect read-only vs state-changing shell commands for agent sandboxes — spammmmmmmmy · 2026-09-28