Adobe's TAC timestamped audio captioning model accepted at NeurIPS, hits SOTA
justin_salamon · x · 2026-09-30
Adobe Research's TAC (Timestamped Audio Captioning) has been accepted at NeurIPS 2026. The model produces timestamped captions for any audio or audiovisual source, tagging overlapping sound events with type labels ([music], [sfx], [speech]).
- TAC-V: fuses TAC outputs with vision-language models for temporally dense audio-visual captions with hallucination correction and visual grounding
- TAC→LLM cascade: acts as a "semantic bridge" for text-only reasoners, achieving SOTA on MMAU-Pro, MMSU, MMAR, Daily-Omni, and VideoHolmes
- Targets the core problem of large audio-language models failing to disentangle overlapping events in complex acoustic scenes
More from Research
- Frontier AI Is a Set, Not a Point: Jagged Capabilities May Be the Steady State — vsikka · 2026-09-30
- UMAP update: new 'recursive' init scales to big data, reproducible multi-core runs — leland_mcinnes · 2026-09-30
- Researcher loses confidence in AA benchmarks, calls them "very misleading" — tianyin_xu · 2026-09-30
- Adelaide's DCSD Decouples Credit Direction and Magnitude in Self-Distillation, Beats GRPO Across 11 Benchmarks — AdelaideUniversity · 2026-09-30
- NVIDIA's HumanoidMimicGen turns one teleop demo into thousands of humanoid demonstrations — AjayMandlekar · 2026-09-30
- Color coding trick yields 2^O(sqrt(n)) depth-3 AC circuits for all symmetric Boolean functions — rrwilliams · 2026-09-30