Deletion Margins, Untouched Histories: AIMS Raises Recall@10 by 4.0-10.9%

When History Misleads: Asymmetric Margin Supervision for Instruction-Guided LLM Generative Recommendation

Ming Yin, Yuhan Yang, Chen Chen, Xinyu Lin, Wentao Shi, Fangcong Yin, Chaofei Yang, Chao Yang, Jiyan Yang, Hui Zhang, Ning Jiang, Yiran Chen, Qifan Wang

cs.IR

2026-10-02

AIMS learns margins from deletions that help the target beat a cutoff rival, training on the full history. Six LLMs, three datasets: Recall@10 beats continued CE by 4.0-10.9%.

What problem this solves

Instruction-guided generative recommenders read a natural-language request and an interaction history at the same time. The request should decide what is acceptable now. History should only personalize among items that already pass. When the two conflict, history often wins. A user who just watched a braised-pork video and then asks for a vegetarian dinner under 30 minutes can see a three-hour beef stew ranked above a 20-minute chickpea pasta.

Knowing which event to remove is the hard part, and a frozen recommender does not answer it. Influence is the Jensen-Shannon divergence between the first item-code distributions with and without that event: a measure of how far two distributions sit, with no sign. On 1,600 industrial training requests and 47,972 events, influence correlates with absolute utility at 0.300 for Llama-3.1-8B and 0.167 for Gemma-3-12B, but only at 0.045 and 0.002 with signed utility. Picking the most influential event beats uniform selection inside the history by 0.24 and 0.42 points in the rate of target-supporting events. Both 95% intervals include zero. Dropping the specific intent also moves the sign: a generic request disagrees with the original 17.10 and 4.65 points more often than a meaning-preserving paraphrase.

A higher target score is not a better rank. On requests that initially miss the top 10, the deletion with the largest target-score gain narrows the margin for 11.8% (Llama) and 16.7% (Gemma), because the item on the cutoff rises more. Re-decoding recovers only 1.08% and 1.50% of those requests as top-10 hits.

Method

Asymmetric Intervention-Guided Margin Supervision (AIMS) uses deletions to build a training target. The student still reads the complete history. Inference is the original decoder, with no deletion search.

Each item is a semantic identifier (SID): three discrete codes derived from content, 2,048 codes per level. A candidate's score is the sum of the three token log-probabilities. The industrial log is searched with SID-constrained beam search, beam 100, over the full catalog. Qilin and KuaiSearch-Lite rank a given candidate pool with the same score.

The frozen supervised-fine-tuning (SFT) checkpoint is both the reference and the student's initialization. The protection set holds training requests whose target is already in the top 10. The competitor stays locked at rank 11. The reference tries every single-event deletion. A deletion is kept only when it raises the target score and the margin over rank 11 by more than 10^-4 nats. The largest accepted margin gain is added to the full-history margin. That sum is the counterfactual margin. If nothing qualifies, the gain is zero and the request stays protected. On Llama and Gemma the protection set is 18.1% and 30.6% of training requests; 88.2% and 85.7% of those have an accepted deletion. Median positive gains are 0.234 and 0.256 nats.

Cross-entropy covers every training request. An auxiliary hinge fires only on the protection set, when the student's full-history margin falls more than 0.01 nats short of the cached margin. The target score is stop-gradient: it decides whether the hinge opens, and the gradient updates only the competitor. Deleting up to three events, or averaging up to five competitors, barely moves Recall@10 and costs up to 2.9 times as many reference forwards. The default stays at one deletion and one competitor.

Results

Six backbones continue from the same SFT checkpoint: Llama-3.1-8B/70B, Gemma-3-4B/12B, and Qwen3-8B/14B. The matched baseline is continued cross-entropy (Continued CE), plus CFT, LETTER (RG), LTRGR, and S-DPO. AIMS is first on Recall@10 and NDCG@10 for every backbone on all three datasets. Relative Recall@10 gains over Continued CE run from 4.0% to 10.9% in every setting. A paired user-cluster bootstrap with 10,000 resamples puts each model ahead of its strongest baseline at p<0.05.

SettingContinued CEAIMS
Industrial, Llama-3.1-8B Recall@106.902%7.223%
Industrial, Gemma-3-12B Recall@107.945%8.609%
NDCG@10, same two models4.460% / 5.141%4.629% / 5.598%
Qilin Recall@10, Llama-8B / 70B35.48% / 43.47%37.61% / 45.62%
KuaiSearch-Lite Recall@10, Qwen3-8B23.94%26.56%

Full-catalog Recall@10 on the industrial log sits near 7% to 9%. Across three seeds, Llama-3.1-8B and Gemma-3-12B beat LETTER by 0.095±0.058 and 0.083±0.027 points, and AIMS is higher in every seed. The 10.9% relative ceiling is Qwen3-8B on KuaiSearch-Lite. Public numbers are pool ranking, so they do not share a scale with full-catalog retrieval.

Ablations on industrial Llama / Gemma: a variant with no deletion gain, only a floor under the original margin, still beats Continued CE by 0.093 and 0.336 Recall@10 points. Those gains account for 71% and 49% of the AIMS lift. Choosing the max-uplift or max-influence deletion leaves an eligible edit on only 75.9/66.1% and 28.8/25.9% of protected requests, against 88.2/85.7% for AIMS, and Recall@10 falls from 7.223/8.609 to 7.091/8.426 and 7.063/8.312. Training on the deleted history with cross-entropy alone does not beat Continued CE. Routing the auxiliary gradient through the target only ties Continued CE.

Retention and recovery both rise. Llama retention moves from 94.73% to 96.28%, recovery from 0.595% to 0.828%. Gemma retention moves from 92.84% to 94.12%, recovery from 0.863% to 1.475%. Recovery supplies 68% and 85% of the Recall@10 gain, even though the auxiliary loss touches only requests that training already ranked correctly. On conflict requests, where at least half of the evaluable history violates the stated category, the Recall@10 gain is 0.558 and 0.874 points, against 0.321 and 0.664 on the full test set. Confirmed violations in the top 10 slots fall from 29.0/25.4% to 21.6/18.3%. The chosen deletion removes a violating event on 80.9/83.8% of the measured protected requests, 16.8 and 16.6 points above the rate implied by the history mix. Under a generic request that excess shrinks to 2.9 and 4.1 points.

Dropping the training history with probability 0.5 trails AIMS by 0.176 and 0.378 Recall@10 points. Swapping in another user's matched history hurts AIMS about as much as Continued CE.

Why it matters

Serving does not change, and there is no deletion search online. The extra cost is offline: the reference scores single-event deletions for the protection set, about 30 views per request, and the cache is then fixed.

The training target lines up with the two measurements above. A per-request margin, with the gradient pushed only through the competitor, is what those measurements ask for. On the industrial full catalog, the three-seed edge of the two reported backbones over LETTER is under a tenth of a point. This is an incremental fix. The SID retrieval stack stays as it was.

Limitations

The cached margin compares the target with one fixed rank-11 item. It is a pairwise stand-in. Top-10 membership still depends on other candidates and on beam search, and the target can enter the list while this margin shrinks.

The conflict slices are observational. Gains rise with conflict severity, which does not identify conflict as the cause. Violation labels cover only industrial test requests that state an explicit positive category, about 25% of that test set. Finer constraints in the pork-rib example, such as a 30-minute cap, are not measured the same way. When the clicked target itself violates the category, selected deletions hit violating events 12.3 and 11.1 points less often than chance. Optimizing the click and obeying the instruction pull apart.

Pools with fewer than 40 candidates are padded with unlabeled training impressions, so public Recall is not on the same scale as full-catalog retrieval. The full-catalog numbers come from anonymized production logs, and the paper describes no public release. The auxiliary loss covers 18% to 31% of training requests on the two ablated backbones. Gains on requests that were initially missed are indirect.

Terms

Source

What people are saying

Related papers

All paper explainers