Phoneme Recognition via Self-Supervised Speech Models

Shikhar Bharadwaj · hf · 2026-07-13

This paper addresses phoneme segmentation and recognition, two closely related tasks that are typically modeled separately.

The authors argue that self-supervised speech models (S3Ms) inherently capture phonological structures and simply need appropriate "guidance." They propose SPAM (Phonological Activation Mapping), which maps frame-level representations to phonological feature activation vectors (e.g., voicing, nasality). This is followed by two lightweight, gradient-free prediction heads dedicated to recognition and segmentation, respectively.

Requiring less than a minute of phonetic annotation, the method generalizes to unseen phonemes and delivers robust performance across multiple datasets.

Original post →

More from Research

Research channel →