Fish Audio S2 fork jumps from 1.2 to 23 t/s on Apple Silicon
fogonthebarrow-downs · reddit · 2026-07-23
A developer forked Fish Audio’s S2 and optimized it for Apple Silicon, claiming a huge speedup on MPS.
- Through quantization and runtime changes, throughput reportedly improved from about 1.2 tokens/s to 12.4 t/s in bf16 and 23 t/s in int8.
- The gains came mostly from KV-cache reuse, GQA expansion, an int8 kernel for MPS, Torch compile on Metal/Inductor, and sub-batch streaming.
- The benchmark was run on an M5 Pro 48GB, and the author is asking other Apple Silicon users to reproduce it on different chips.
- This is a useful local TTS optimization post for people working on Apple-device inference.
More from Multimodal
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11