Speech Model Shrunk 13x to 153M Params by Looping 2 Shared Blocks

pbaylies · x · 2026-10-11

Oruk Labs compressed a speech model 13x by making it "repeat itself": 52 layers were replaced with 2 shared blocks that loop—encoder loops 24 times, decoder 28 times—cutting parameters from 2.04B to 153M while keeping compute roughly the same.

This recurrent-depth approach trades weights for serial computation steps, a promising recipe for on-device speech models.

Original post →

More from Infra

Infra channel →