GPT-4o hit 70% on spectrogram classification two years ago, near human experts' 72.5%
TuhinChakr · x · 2026-09-10
Chris Donahue points out that VLM-based audio spectrogram classification isn't new: in the paper "Vision Language Models Are Few-Shot Audio Spectrogram Classifiers," GPT-4o scored 70% on 10-way spectrogram classification versus human experts at 72.5% — something possible for at least a couple of years. He jokes he's entering the "we tried that years ago" phase of his research career.
Context: Greg Brockman was showing Astra doing spectrogram-to-sound identification.
More from Models
- OpenAI developers site adds showcase page with GPT-6 Astra outputs — Dimillian · 2026-09-10
- Arena.ai: Claude Fable 5.1 Writes More Matter-of-Fact but More Verbose — The Decoder · 2026-09-10
- Qwen-Image-Edit-2511 vs SenseNova U1.5 Lite: hands-on multi-reference fusion comparison — daniel933912 · 2026-09-10
- Cognition launches SWE-2: frontier-level coding performance at up to 70% lower cost — silasalberti · 2026-09-10
- NeoHorse-1-4B, a Qwen3.5-based agentic model, trends on Hugging Face — TokenRhythm · 2026-09-10
- Chinese model 3D face-off: DeepSeek V4.1 Flash crushes Kimi K3 and GLM-5.3 — teortaxesTex · 2026-09-10