Speech Model Shrunk 13x to 153M Params by Looping 2 Shared Blocks
pbaylies · x · 2026-10-11
Oruk Labs compressed a speech model 13x by making it "repeat itself": 52 layers were replaced with 2 shared blocks that loop—encoder loops 24 times, decoder 28 times—cutting parameters from 2.04B to 153M while keeping compute roughly the same.
This recurrent-depth approach trades weights for serial computation steps, a promising recipe for on-device speech models.
More from Infra
- Inference demand went vertical, yet is a flat line next to post-training/RL growth — zainhas · 2026-10-11
- Pat Gelsinger slams HBM as "a lousy memory" wasting four bits for every one it makes — SumitGup · 2026-10-11
- Qualcomm CEO predicts AI phone supercycle, smart glasses as top AI wearable — SuB8u · 2026-10-11
- Zero cold starts: Building and shipping MCP servers with WebAssembly, Spin and Akamai Functions — AI Engineer · 2026-10-11
- Hugging Face launches a PyTorch profiling series: from torch.profiler to attention — ariG23498 · 2026-10-11
- Bain sees 183GW of new data center capacity by 2030, needing $5-6.5T in spending — Beth_Kindig · 2026-10-11