Phonon-2 hits 606x real-time on a MacBook Air: one hour of speech in 6 seconds
julianweisser · x · 2026-10-07
Manan announces a Core ML package for the speech model Phonon-2, reaching 606x real time on a base M5 MacBook Air — up from 174x at launch — transcribing an hour of speech in about 6 seconds with no quality loss. It's now the default engine in the app Detta, with peak memory cut from 3.1GB to 1.1GB and idle memory from 1.8GB to 600MB even with Gluon loaded, matching WisprFlow's footprint while running fully locally.
More from Infra
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- HF engineer releases open slide deck on local AI: quantization to speculative decoding — mervenoyann · 2026-10-07
- 21M model + 6.4B SSD-resident lookup table matches a 114M dense model — fechyyy · 2026-10-07
- Ollama 0.35 adds local decision models from Cloudflare, Together and Bespoke — Technovangelist · 2026-10-07
- CostGraph becomes a drop-in InfraCost replacement for comparing GPU prices from L40 to A100 — saheedniyi_02 · 2026-10-07
- Oki Home launches a $1,799 Memory Computer running a 27B model locally with up to 16TB Memchip storage — Scobleizer · 2026-10-07