EnCodec at 3 kbps Beats Lyra-v2 at 6 kbps in MUSHRA Tests

High Fidelity Neural Audio Compression

Alexandre Défossez, Jade Copet, Gabriel Synnaeve, Yossi Adi

eess.AS, cs.AI, cs.SD, stat.ML

2022-10-25

EnCodec is a streaming neural codec with RVQ and one STFT discriminator. At 24 kHz, 3 kbps beats Lyra-v2 at 6 kbps on MUSHRA; 48 kHz stereo at 6 kbps matches MP3 at 64 kbps.

What problem this solves

Classical codecs such as Opus and EVS fall apart at very low bitrates. Neural codecs compress harder, but training often stacks several waveform discriminators whose loss scales fight each other, and streaming adds padding and normalization issues. SoundStream / Lyra already showed that residual vector quantization can serve many bitrates with one model. The stack is still heavy.

EnCodec aims at a real-time, streamable neural codec that holds up from speech to music: 24 kHz mono from 1.5 to 24 kbps, plus a 48 kHz stereo music setting. One set of weights should cover multiple rates and run faster than real time on a single CPU core.

Method

The network is a convolutional encoder, residual vector quantization, and a mirrored decoder. The encoder starts at 32 channels, uses four downsampling stages with strides (2, 4, 5, 8), residual blocks, and a two-layer LSTM. At 24 kHz it emits 75 latent frames per second; at 48 kHz, 150. The streamable variant puts all padding on the left, so 320 samples in (about 13 ms) yield 320 samples out, and swaps time-pooled layer norm for weight norm.

Quantization is RVQ: up to 32 codebooks (16 at 48 kHz), 1024 entries and 10 bits each. Training keeps a random multiple of four codebooks so one model covers 1.5, 3, 6, 12, and 24 kbps. Codebooks update by EMA, dead entries are replaced from the batch, and a commitment loss is added.

The training objective mixes time-domain L1, multiscale mel L1+L2, and a single multiscale STFT discriminator (five windows, real and imaginary STFT stacked). A relative feature-matching term normalizes by discriminator-layer energy. The stabilizer is a loss balancer: each weight is the share of the generator gradient that loss should contribute, independent of the raw loss scale. Discriminators are per-bandwidth; a batch updates only the one that matches the sampled rate.

A 5-layer Transformer language model plus arithmetic coding trims another 25% to 40% of bandwidth, at the cost of waiting one extra frame (about 13 ms).

Results

24 kHz streamable MUSHRA (higher is better):

ModelRateClean speechNoisy speechMusic set 1
EnCodec3 kbps (1.9 after entropy)67.062.589.6
EnCodec6 kbps (4.1)83.169.492.9
Lyra-v26 kbps66.259.975.7
Opus12 kbps76.561.977.8
EVS9.6 kbps84.480.089.9

At matched nominal rates EnCodec leads every baseline. Average quality at 3 kbps already beats Lyra-v2 at 6 kbps and Opus at 12 kbps. On music, 3 kbps is close to EVS at 9.6 kbps.

48 kHz stereo music: 6 kbps scores 82.9 MUSHRA, in line with MP3 at 64 kbps (82.7) and Opus at 24 kbps (82.9); 12 kbps reaches 88.0. Entropy coding saves about 20% to 30% more.

Ablations: MS-STFT alone yields MUSHRA 77.5, SI-SNR 6.67, ViSQOL 4.35. Adding MPD adds 1.5 MUSHRA points and more training cost. Streamable versus non-streamable at 6 kbps: SI-SNR 6.67 vs 7.46, ViSQOL 4.35 vs 4.39. On a 2019 MacBook Pro, one thread, 6 kbps: 24 kHz encode RTF 9.8, decode 10.4; with entropy coding still about 1.6. The 48 kHz non-streamable model has a 1 s initial latency and drops below RTF 1 with entropy coding.

Why it matters

This is a parent of later open audio tokenizers. Mimi and a line of speech LMs sit on the same RVQ streaming-codec idea. Three copyable choices: a spectrogram-only discriminator is enough; set loss weights as gradient shares rather than raw magnitudes; a tiny LM plus arithmetic coding buys extra compression when latency is not the bottleneck.

Beating 6 to 12 kbps baselines at 3 kbps is a listening-test gap, not a one-point table win.

Limitations

MUSHRA is crowdsourced on fifty 5-second clips with at least ten kept ratings each. Short clips do not prove long-form stability. 48 kHz is trained and tested on music only. Streaming has a measurable objective drop. Entropy coding needs matched floats across machines; the authors round probabilities to 1e-6 and flag this as unfinished for production. Lyra-v2 is run on audio upsampled to 32 kHz, so sample rates are not aligned. The small Transformer yields less extra compression at high rates, which they attribute to modeling many codebooks at once.

Terms

Source

What people are saying

Related papers

All paper explainers