VibeVoice 1.5B Runs Locally on iPhone: ~2.2GB RAM and 1.28x Real-Time Speed
Acceptable-Cycle4645 · reddit · 2026-08-05
A developer has successfully achieved local inference of the VibeVoice 1.5B model on an iPhone. This edge deployment test yielded impressive efficiency metrics:
- Low memory footprint: Requires only about 2.2GB of memory during operation.
- Generation speed breakthrough: Achieves up to 1.28x real-time speed for long-form generation, with stable VRAM/memory usage.
The author noted that the model weights have been uploaded to the audio.cpp Hugging Face repo. The xcframework and core code will be pushed to a branch after the 0.6 release, offering a new paradigm for local mobile audio inference.
Related event: VibeVoice 1.5B Runs Locally on iPhone(2 posts)→
More from Infra
- Astera Labs Predicts NPO Deployment in 2027, CPO to Follow in 2028 — bookwormengr · 2026-08-05
- US Chip Export Controls Backfire: Samsung and SK Hynix Turn to Chinese Toolmakers — kevinsxu · 2026-08-05
- TriAttention Integrated into TensorRT-LLM for Efficient Long-Context Inference — songhan_mit · 2026-08-05
- SpaceX Market Cap Drops $130B Overnight After $15.8B AI Spending Spree in Q2 — 智东西 · 2026-08-05
- Ex-OpenAI Exec Slams Goldman Sachs Token Demand Forecast, Cites 100x Cost Drop — ChrSzegedy · 2026-08-05
- SK Hynix and Samsung Evaluate AMEC Etchers for Chinese Fabs — zephyr_z9 · 2026-08-05