Qwen 3.8 on Apple Silicon speeds up nearly 3x using AI-written kernels
alexcovo_eth · x · 2026-08-18
Qwen 3.8 27B is now running 193.4% faster (2.94x) on Apple Silicon compared to last Friday. The acceleration was achieved by using AI models like DeepSeek V4 Pro, GPT-5.6, and Claude to analyze an open leaderboard and rewrite/optimize the MLX kernels automatically.
More from Infra
- Model Prices Collapse, but Gateway Fees Remain High — AccBalanced · 2026-08-18
- Performance Analysis of Mixing RTX 3080 and A2000 for Stable Diffusion — djekler · 2026-08-18
- Snapdragon X2 Elite Extreme Doubles AI Performance Over Rivals — ryanshrout · 2026-08-18
- Salvaging parts: building a local AI rig with 64GB RAM and a €1700 budget — joquinjack · 2026-08-18
- Mesh LLM: A Third Option Between Crypto Rigs and Cloud Subscriptions — alex_verem · 2026-08-18
- Overclocking VRAM on 4x RTX 5060Ti for LLM inference — Ok-Breakfast1878 · 2026-08-18