YuE2 music model with full CoT runs on iPhone in just 1.7GB of memory
Acceptable-Cycle4645 · reddit · 2026-09-25
Developer 0xShug0 demoed running YuE2 (ABC codec + full chain-of-thought) on an iPhone with only 1.7GB of memory, pushing the music generation model's VRAM usage to an extreme for on-device inference. The optimization code is open-sourced in the yue2-optimizations branch of the audio.cpp repo on GitHub.
More from Infra
- VeriTile embeds Triton GPU kernels in Lean, with AI agents writing machine-checked correctness proofs — KaiyuYang4 · 2026-09-25
- Merge Gateway Launches Batch Inference at ~50% of Standard Prices — shensi · 2026-09-25
- New deep-dive article on scaling LLM inference in production — abhijithneil · 2026-09-25
- Lambda engineer shares local inference build rule: 27B models need 24-32GB VRAM — TheZachMueller · 2026-09-25
- Pokee AI demos 36B agent model running fully local on Snapdragon X2 Elite with 32GB RAM — Kyrannio · 2026-09-25
- AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026 — PyTorch · 2026-09-25