kalomaze: coarse audio features that tell you ~nothing still suffice for gender classification
kalomaze · x · 2026-09-21
kalomaze points out an asymmetry in speech modeling: a bundle of extremely coarse audio features carries no information about what was spoken, the speaker's age, language, identity, or emotion — yet is sufficient for gender classification.
His argument: when such asymmetry exists, how much a model learns about the conditioning depends on how much predicting the labels well forces it to model the conditioning. BCE gender classification of human speech is an extreme case, where the model needn't learn the rest of speech structure at all.
Related event: Developer Says Multimodal Training Still Needs SSL Backbones(4 posts)→
More from Research
- PPO was rejected from NeurIPS in 2017; now it's the most-used RL algorithm with 45,000+ citations — IanArawjo · 2026-09-21
- Jev, Prolog, Pi, and the dream of probabilistic logic programming — schmuhblaster_x45 · 2026-09-21
- Chemist Critiques Biology Millennium Problems: All Engineering, No Discovery or Mechanism — burny_tech · 2026-09-21
- Model Sweeps Every Deposited Ribosome Structure, Finds 240 Folding in Ways No Force Field Allows — JosephJacks_ · 2026-09-21
- FutureHouse Publishes 'Millennium Problems for Biology' as the Ultimate Eval for AI in Bio — burny_tech · 2026-09-21
- ModularRSI Paper Names Three Defects in How Agent Harnesses Get Improved — and the Fixes — dair_ai · 2026-09-21