Fermion open-sources Phonon-2: a 164MB speech model matching 2.5GB teachers, 1h audio in 20s on MacBook Air
solyarisoftware · x · 2026-09-30
Fermion Research released Phonon-2, an open English ASR model it calls the most accurate open speech recognizer under 900MB.
- Tiny download, high accuracy: a 164MB download averaging 5.21% WER across the Open ASR Leaderboard's seven English sets; every open model scoring better is at least 5.8x larger
- Quantization: encoder stores each weight as one of five learned levels at 2.1 bits, matching the accuracy of its 2.5GB full-precision teacher (derived from NVIDIA Parakeet TDT 0.6B v3, CC-BY-4.0) and beating it on meetings and parliamentary speech
- Noise robustness: leads Parakeet Redux at every noise level tested
- Fast: 20 seconds to transcribe an hour of audio on a MacBook Air; runs on Mac, Linux, Windows and NVIDIA GPUs
- Ships via pip, CPU/NVIDIA Docker images, HF weights, plus Detta, a Mac dictation app
Related event: Open-Source Phonon-2 ASR Model Beats Whisper Large at Just 164MB(3 posts)→
More from Models
- Anthropic retires Claude Opus 3 but keeps it on API and gives it an essay column — repligate · 2026-10-01
- Bindu Reddy: Gemini Argon pricing is 5x cheaper than Astra, but benchmarks look too good — bindureddy · 2026-10-01
- Gemini Answers Niche Questions Claude Can't, Says User Pushing Back on Programmer Gripes — PAstynome · 2026-10-01
- User finds Opus 5.5 Max still reproduces Opus 5's broken outputs — 0xkarasy · 2026-10-01
- Gemini 4 impresses on benchmarks, says a longtime Google model critic — iruletheworldmo · 2026-10-01
- Fable 5.1 roleplay instance plans 5-turn solo quest and rarely addresses users by name — repligate · 2026-10-01