No fine-tuning needed: injecting audio vectors into KV cache turns an LLM into an audio model

tkasasagi · x · 2026-09-28

A new arXiv paper (submitted to ICASSP) proposes a symbiotic architecture where an injector module writes audio-conditioned vectors directly into a frozen LLM's KV cache, making it behave as an audio language model without any weight updates. Injection cost scales with injector width rather than backbone width; the approach beats standard frozen-LLM baselines and approaches fine-tuned ALMs on ASR, audio QA, and scene classification while preserving text performance by construction.

Original post →

More from Research

Research channel →