Sona replaces Yandex Music's 15+ generator cascade, lifting active users 4.53%

Sona Technical Report

Sona Team, Alexandr Udeneev, Aleksei Krasilnikov, Alexey Nadtochiy, Andrey Semenov, Andrey Tsyrkunov, Anna Krivonos, Anna Lipkina, Artem Matveev, Daniil Burlakov, Daniil Leshchev, Daria Tikhonovich, Denis Burshtein, Ekaterina Dmitrieva, Eugene Krofto, Grigorii Khlystov, Ilya Murzin, Kirill Golovko, Ksenia Sycheva, Leonid Dmitriev, Mariia Rozaeva, Mariia Ulianova, Mikhail Sandul, Nikolai Savushkin, Oleg Sorokin, Roman Odobesku, Semyon Panenko, Sergei Liamaev, Sergei Makeev, Vadim Shilov, Veronika Ivanova, Viktor Yanush, Vladimir Baikalov, Vladislav Dodonov, Vladislav Tytskiy

cs.IR

2026-08-11

On smart-speaker My Vibe, Sona replaces Yandex Music's cascade of 15+ generators plus rankers, lifting active users 4.53%, listening time 6.30%, and likes 11.42%.

What problem this solves

Music feedback is thin. Tracks play in the background, so the log is mostly plays and skips, with a like only now and then. Hearing a familiar song again is normal, and that repeat is the signal. Familiar tracks hold listening time now. A new favorite pays off later. The ranker has to carry both.

My Vibe on Yandex Music smart speakers is that job with the user intent removed. Playback starts with no artist, genre, or mood, and the surface is one of the largest by listening time. Production was a cascade of more than 15 candidate generators, then pre-ranking and ranking. The ranker consumed hundreds of features, including large transformers such as Argus and target-attention scorers. Gains from V0, V1, V2, and Argus were already in the control. Sona has to replace the whole cascade and still add lift on top.

Method

Generation and ranking share one encoder pass. The encoder reads the chronological stream of plays, skips, likes, dislikes, track identity, and request surface, and both the decoder and the Ranking Module read that state. A request does not encode the user twice.

Tracks are emitted as Semantic IDs, not raw item ids: 3 codes per track, three codebooks of 32,000. A frozen Qwen2.5-Omni runs prefill only, on the mel spectrogram of the first 90 seconds plus title, artist, and tags. A 4-layer, 4-head, width-512 network refines that vector on collaborative pairs, mean-pools it to 256 dimensions, and residual K-means quantizes it coarse to fine. The pairs are disjoint on purpose. Cross-artist tracks shown together in one production response contribute about 220 to 240 million pairs over three weeks. Same-artist, high-NPMI co-listens contribute about 20 million pairs over three months. Neighbors in audio space are not always tracks listeners treat as substitutes, so the loss is InfoNCE plus an alignment term weighted 0.1.

Serving uses trie-constrained beam search at width 1024. Codes collide, so each finished triple expands to the tracks that share it, and the ranker orders that set. Candidate features are the item id, the codes, and prefix n-grams. Cross-attention against the shared user state produces one head per teacher score, fit by mean absolute error. The teacher is absent at serving. A fixed weighted sum collapses the heads, and that scalar is outside the training loss.

The teacher is larger. It attends the full history with a causal encoder, skips History Compression, and trains on a year of logs. The student's joint run covers several weeks, then online updates continue. The year is compressed into the teacher because every student sample re-encodes the whole history to supervise a few new targets. The teacher first does next-item contrastive pre-training, then pairwise ranking on the grade order like, play, skip, dislike.

The student loss is next-token prediction on engaged impressions, plus distillation on decoder rollouts and on logged impressions. The training beam is 32, so the teacher signal follows the candidates the decoder currently emits, and the gradient updates the shared encoder.

History is capped at 8,192 events. The deep stack of 7 layers runs only on the latest 2,048. The older 6,144 stay on a shallow path, and downstream modules still see every position, at about half the inference cost of a full encoder. Continuous values are bucketed. There are no hand-built cross features.

Weights ship every 10 minutes. Sessions wait at least 15 minutes so the inference log can catch up. From an emitted event to a production model trained on it, median delay is 45 minutes and the 99th percentile is 60. Each training sample is rebuilt at the history cutoff stored at serving time, so the gradient matches the state that produced the candidates.

Results

The tokenizer gain is in the representation, not the codebook size. With 3×32,000 codes, Qwen2.5-Omni plus collaborative refinement reaches Target-track Recall@10/100/1000 of 0.2171 / 0.5362 / 0.8524. CLMR audio embeddings at the same codebook size sit at 0.1848 / 0.4712 / 0.8111. Growing the CLMR codebook from 3×8,192 to 3×32,000 only moves Recall@1000 from 0.8036 to 0.8111.

On a Medium backbone, 2,048-event histories, and a single pass, stretching the log from one week to eight weeks lifts Recall@10 from 0.1798 to 0.2519 and Recall@1000 from 0.8333 to 0.8947. Core sizes of 20M, 130M, and 260M, or 264M, 485M, and 659M once embeddings and heads are counted, keep lowering training next-token loss over the measured exposure.

Pre-training barely moves the teacher offline. At 8,192 events and a 6-layer scorer, weighted pair accuracy goes from 0.6153 to 0.6215. Eight layers at that history length reach 0.6219. Online, the full two-stage teacher replaces the production ranker on 3% of users: listening time +1.64%, active users +1.22%. Likes at +1.48% miss p < 0.01.

Decoder likelihood does not preserve the teacher's order. Teacher Recall@10 is 0.0381 with no Ranking Module, 0.5654 with one scorer layer, and 0.6005 with four, on 2,048 events with 7 encoder layers and 2 decoder layers. Full attention at 8,192 events and 4 scorer layers reaches Target-track Recall@1000 of 0.8722 and Teacher Recall@10 of 0.6586. History Compression lands at 0.8709 and 0.6474, with WPA 0.6033 against 0.6029. The shipped model uses compression.

MAE, MSE, Huber, and pairwise KL: none wins every metric, so later runs keep MAE. Distilling on logged impressions alone drops Teacher Recall@10 to 0.2983. Rollouts plus impressions hold 0.5654. A training beam of 128 scores 0.5762, and a beam of 32 scores 0.5654. They keep 32.

In experiment 4 the distilled model, still on 2,048-event histories, replaces the full cascade: active users +1.41%, listening time +1.62%, likes +7.12%. The same generator re-ranked by the teacher reaches +2.81%, +4.32%, and +11.81%. The single model can ship. At this history length it is still well behind the teacher.

Experiment 5 is the full recipe: 8,192 events with History Compression, 15% of users per arm, seven days, and the production business rules left in place. Every delta below is significant at p < 0.01.

Metricvs production control
Active Users (primary)+4.53%
Total Listening Time+6.30%
Likes+11.42%
Repeat Commands+17.99%
Deeply Engaged Users+7.37%

The +4.53% on active users is 2.35× the +1.93% Argus previously added on this surface, on top of gains already kept from V0 through Argus. Experiments 4 and 5 were separate tests. Their gap is not a clean read on history length.

Why it matters

A production cascade is many retrievers plus a ranker stuffed with features. Sona turns that into one encode, one short-code decode, and one score against the same state. The teacher stays off the request path.

On a passive, long-session product, the short codes let a new track enter generation through a shared prefix before it has a behavior history. Weights update every 10 minutes, and the median loop from event to serving is 45 minutes. Inference MFU is 41% at beam 1024. Experiment 5 kept the hard business rules, so the lift is not from dropping constraints. The control already included Argus. This is an increment on a mature stack.

Limitations

Full traffic is still off. The report wants a multi-month test first, and it says catalog coverage is lower than production, cause unknown. Only smart-speaker My Vibe was measured. Search and playlists were not. There is no RL, and test-time compute is discussed only as beam width.

Exploration was not measured. The report sets up a real tension, familiar tracks for time spent and new tracks for a longer horizon. Experiment 5 raises Repeat Commands 17.99% while catalog coverage falls. It says exploration did not look worse, and points at training on production logs, but it gives no fresh-track share. The gains in time and likes fit a model that is simply better at familiar music. Those stories are not separated.

Pre-training is described as the stage that lets a feature-free teacher replace a ranker with hundreds of features. Offline, WPA only moves from 0.6153 to 0.6215, and there is no online ablation of that stage. The jump from +1.41% active users in experiment 4 to +4.53% in experiment 5 mixes history length with other recipe changes across two separate tests. The report says this is not an isolated causal effect.

Serving ships one model. Training still runs a teacher that reads a year of full history and refreshes daily, while the student is distilled on a much shorter window.

Terms

Source

What people are saying

Related papers

All paper explainers