SonicCaps: 15M-caption audio dataset improves CLAP retrieval and zero-shot classification
serrjoa · x · 2026-09-03
- SonicCaps is a large-scale audio captioning dataset: 15M captions over 700k audio clips, generated with Qwen3-Omni conditioned on audio and text.
- Diversity is engineered via structured prompts and few-shot generation — 24 captions per audio spanning main descriptions, rephrased variants, and semantic tags, addressing the low diversity and one-to-one mapping limits of existing datasets.
- Human evaluation rates the captions significantly more descriptive and precise; training CLAP with multi-caption sampling consistently improves audio retrieval and zero-shot classification. Dataset and two CLAP models released on Hugging Face.
More from Multimodal
- ComfyUI plugin caches MiniMax H3 conditioning to disk: 1.12s cache hits, 14GB RAM saved — Any_Fee5299 · 2026-09-03
- Open-source ArtSmoker pipeline turns text prompts into textured, Blender-ready 3D models on your own AWS — niravdd · 2026-09-03
- Midjourney character sheet + Seedance long video recreate a 70s Giallo thriller — Kyrannio · 2026-09-03
- Catwoman character explorations generated with Midjourney — hewarsaber · 2026-09-03
- Fixing AI anime background inconsistency by pre-shooting a 360° orbit video with MiniMax H3 — Hailuo_AI · 2026-09-03
- Fable 5.1 one-shots a full quad-turbo W16 engine CAD model and animation — repligate · 2026-09-03