AudioRubrics: Enhancing AI Audio Reasoning with Self-Evolving Reward Mechanisms

UMCP · hf · 2026-08-10

Existing RL-based audio reasoning models face reward limitations: outcome-based rewards ignore the reasoning process, while process-based rewards rely on rigid, hand-crafted criteria. To address this, researchers introduce AudioRubrics, a framework that supervises audio reasoning with self-evolving, audio-grounded rubric rewards.

The framework synthesizes per-sample rubrics from raw waveforms and dynamically adjusts criteria based on the model's own rollouts, continuously targeting current policy weaknesses. Across three audio reasoning benchmarks, AudioRubrics significantly outperforms various open-source baselines. Analysis shows its gains scale with rubric generator capability, and it converges to a stable reasoning length, avoiding both degenerate collapse and unbounded growth.

Original post →

More from Research

Research channel →