SpeechLMs secretly transcribe: implicit text-decodable stage found in middle layers
kastnerkyle · x · 2026-09-10
- Key finding: interleaved Speech LMs enter an implicit transcription stage — spoken words become text-decodable in middle layers without ASR supervision, suggesting speech models may "think" in text.
- Authors pose: is this a feature, shortcut, or limitation?
- Accepted to EMNLP 2026 Findings; the updated version adds natural-speech evaluation, additional decoding methods, and new controls.
More from Research
- Pretraining study: varied auxiliary views beat repetition for LLM knowledge acquisition — kastnerkyle · 2026-09-10
- ByteDance Seed Unveils ByteWrist: A Parallel Robotic Wrist for Confined-Space Manipulation — scott_e_reed · 2026-09-10
- The Prism Hypothesis unified autoencoding paper accepted at ECCV 2026, paves way for encoder-free MLLMs — liuziwei7 · 2026-09-10
- embedflow Migrates Embedding Models Without Re-embedding: Qwen 4B→8B Matches Native Retrieval with Just 50 Reranked Docs — Potential_Low_1183 · 2026-09-10
- Degenerate Fisher information explains why huge neural nets don't defy Occam's razor — FrnkNlsn · 2026-09-10
- AIxBio researcher: skip the bitter lesson debate — more compute means lower per-dollar efficiency — anshulkundaje · 2026-09-10