kalomaze: coarse audio features that tell you ~nothing still suffice for gender classification

kalomaze · x · 2026-09-21

kalomaze points out an asymmetry in speech modeling: a bundle of extremely coarse audio features carries no information about what was spoken, the speaker's age, language, identity, or emotion — yet is sufficient for gender classification.

His argument: when such asymmetry exists, how much a model learns about the conditioning depends on how much predicting the labels well forces it to model the conditioning. BCE gender classification of human speech is an extreme case, where the model needn't learn the rest of speech structure at all.

Related event: Developer Says Multimodal Training Still Needs SSL Backbones(4 posts)→

Original post →

More from Research

Research channel →