Audio2Face-3D Ported to Apple Silicon for Local 3D Animation

ivan_digital · reddit · 2026-07-05

Developers have ported NVIDIA's Audio2Face-3D (audio-driven 3D facial animation) model to run locally on Apple Silicon as part of an open-source (Apache 2.0) speech-swift package. The forward pass uses a handwritten MLX compute graph, eliminating the need for an ONNX runtime. It takes a WAV input and outputs timestamped facial animation coefficients with emotion adjustment support, releasing three MLX models: James, Claire, and Mark. The package also provides local TTS and voice cloning, enabling an offline 'text -> cloned voice -> facial motion' pipeline on Mac.

Original post →

More from Multimodal

Multimodal channel →