Audio2Face-3D Ported to Apple Silicon for Local 3D Animation
ivan_digital · reddit · 2026-07-05
Developers have ported NVIDIA's Audio2Face-3D (audio-driven 3D facial animation) model to run locally on Apple Silicon as part of an open-source (Apache 2.0) speech-swift package. The forward pass uses a handwritten MLX compute graph, eliminating the need for an ONNX runtime. It takes a WAV input and outputs timestamped facial animation coefficients with emotion adjustment support, releasing three MLX models: James, Claire, and Mark. The package also provides local TTS and voice cloning, enabling an offline 'text -> cloned voice -> facial motion' pipeline on Mac.
More from Multimodal
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11