Testing Gemini as music understanders: Pro 3.1 solid, Flash models hallucinate sounds
teropa · x · 2026-09-07
teropa tested Gemini models as music understanders. Pro 3.1 reasonably identifies instrumentation, style and vocal presence. Flash 2.5 and 3.8 produce many false positives, imagining synth pads everywhere (2.5 also adds phantom drums). Increasing thinking makes Flash worse, giving more room to confabulate — suggesting the limits lie in perception, possibly audio tokenization.
More from Models
- Blogger revises AI training scale estimate to 5-7T, says 8T already too generous — scaling01 · 2026-09-07
- OpenAI's Astra math 'breakthroughs' commit research misconduct, mathematicians say — asusarla · 2026-09-07
- Peter Gostev debunks model sparsity leak: Kimi 26:1, DeepSeek 32:1, 1.2T active params implausible — inductionheads · 2026-09-07
- OpenAI's newest models block function tools on /v1/chat/completions, forcing Responses API migration — AI-Specialist-6597 · 2026-09-07
- Astra excels at verifier-backed goals but lags Sol at instruction following, dev observes — MinqiJiang · 2026-09-07
- Anthropic, Google, and OpenAI's $1 federal government contracts expire this month — LuizaJarovsky · 2026-09-07