Apple details latent-space distillation to shrink on-device streaming audio encoders

Apple ML Research · rss · 2026-09-24

Apple ML Research published 'Compressing Streaming Neural Audio Encoders via Latent-Space Distillation', covering the always-on tokenizer behind fully on-device system-wide dictation. With the foundation model sparsely activated via Instruction-Following Pruning, only a subset of experts occupies DRAM, so the tokenizer competes for memory and its size drives power and latency. The work distills the encoder against latent-space supervision to compress it.

Original post →

More from Infra

Infra channel →