oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads
awnihannun · x · 2026-07-22
oMLX 0.5.2 ships a batch of Mac-focused runtime improvements for running MLX models locally.
Highlights
- New menu bar live stats for throughput and CPU/GPU/MEM usage, with popovers and rolling graphs
- A reorganized Models menu that groups Loaded, Favorites, and Library
- Bonsai 1-bit / 2-bit decode kernels for extreme low-bit models
- Faster Hugging Face downloads, more TTS output formats, and chat history import/export
- Additional cache, memory, and stability fixes for long-running servers
The release also says earlier 0.5.x work improved prefill throughput on several models, including DeepSeek-V4-Flash and Qwen3.6-27B.
More from coding & agent
- Gemini 3.5 Flash-Lite is 71x Cheaper Than Claude for Doc Extraction — rseroter · 2026-07-23
- LangChain and Cognition will host a meetup on open memory for agents — LangChain · 2026-07-23
- Factory Co-founder Predicts 90% of Coding Agent Tokens Will Be Fully Autonomous in 12-24 Months — matanSF · 2026-07-23
- Voice assistant tool calls sped up instantly after moving the backend to Europe — ur_piyo_a_hoe · 2026-07-23
- W&B’s Scott Condron wants to push research agents, trace insights, and marimo eval UIs — _ScottCondron · 2026-07-23
- Clean Plate LoRA examples show a practical video-cleanup workflow — nazihater3000 · 2026-07-23