FastH3 Now Runs Locally on Apple Silicon and DGX Spark
Vandy_simp · reddit · 2026-09-03
The FastVideo team released local inference paths for FastH3: via MLX on Apple Silicon and on one or two NVIDIA DGX Sparks. Maintained recipes are exposed through Python, CLI tools, and a local OpenAI-compatible server/playground, covering setup, weight conversion, memory constraints, and reproducible generation. The team notes it is not yet a drag-and-drop ComfyUI workflow and is asking users whether native nodes, an API-backed node, or example workflows would be most useful. Code is open-sourced on GitHub.
More from Infra
- llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default — Bulky-Priority6824 · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing — The Verge AI · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03