No fine-tuning needed: injecting audio vectors into KV cache turns an LLM into an audio model
tkasasagi · x · 2026-09-28
A new arXiv paper (submitted to ICASSP) proposes a symbiotic architecture where an injector module writes audio-conditioned vectors directly into a frozen LLM's KV cache, making it behave as an audio language model without any weight updates. Injection cost scales with injector width rather than backbone width; the approach beats standard frozen-LLM baselines and approaches fine-tuned ALMs on ASR, audio QA, and scene classification while preserving text performance by construction.
More from Research
- Researcher Argues Self-Supervised Learning Is Underhyped, Partly Rebranded as World Models — iScienceLuvr · 2026-09-28
- Quadruped Gait Atlas 3D: Muscle-Driven Skeletons Solve Lowest-Cost Gaits at Every Speed — burny_tech · 2026-09-28
- MICCAI Researcher: Recent NeurIPS Spectral Learning Papers Ignore the Strasbourg Gamma-Calculus Tradition — PTenigma · 2026-09-28
- AnyMo, a Setup-Agnostic IMU Human Motion Model, Accepted at NeurIPS 2026 — flosalim · 2026-09-28
- MIT's Functional Differential Geometry book is now free via Open Access — FrnkNlsn · 2026-09-28
- GPT-6 trains a tiny 12k-param CNN to play VizDoom at 35 fps realtime — paraschopra · 2026-09-28