Audio2Face-3D Ported to Apple Silicon for Local 3D Animation
ivan_digital · reddit · 2026-07-05
Developers have ported NVIDIA's Audio2Face-3D (audio-driven 3D facial animation) model to run locally on Apple Silicon as part of an open-source (Apache 2.0) speech-swift package. The forward pass uses a handwritten MLX compute graph, eliminating the need for an ONNX runtime. It takes a WAV input and outputs timestamped facial animation coefficients with emotion adjustment support, releasing three MLX models: James, Claire, and Mark. The package also provides local TTS and voice cloning, enabling an offline 'text -> cloned voice -> facial motion' pipeline on Mac.
More from Multimodal
- Midjourney V8.2 adds personalization and shows off stylized image outputs — Mr_AllenT · 2026-07-27
- Midjourney’s image variety draws a Krea 2 comparison and asks how to reproduce it — diffusion_throwaway · 2026-07-27
- AI short film sets a 1985 dystopia to music and leans into cinema — ProfessorKey98 · 2026-07-27
- A new BOTPD episode made with Google Omni turns into an AI chase-scene parody — ScriptLurker · 2026-07-27
- A new LoRA recreates GTA: San Andreas’ classic RenderWare-era visuals — Humble-Pick7172 · 2026-07-27
- Enabling dynamic VRAM cuts LTX 2.3 video generation to 168s on an AMD R9700 — xdcfret1 · 2026-07-27