Huawei's GrIS Reframes Semantic ID Construction as Graph Partitioning, Gaining up to 52% Hit@10

Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)

Aleksei Medvedev, Alejandro Ariza-Casabona, Steven Derby, Gonzalo Fiz Pontiveros, Xinyang Shao, Florian Spiess

cs.AI, cs.IR, cs.LG

2026-10-01

Huawei's GrIS casts Semantic ID construction as hierarchical partitioning of an item graph, subsuming RQ-VAE as the empty-graph case; Hit@10 gains reach +52% over LETTER.

What problem this solves

Generative recommendation drops nearest-neighbour retrieval: a model autoregressively generates the identifier of the next item. Those identifiers are Semantic IDs (SIDs), short sequences of discrete codes per item. The dominant construction treats this as representation learning: encode item content into a vector, quantise it with RQ-VAE-style residual quantisation, read off the codes. TIGER is the reference point.

The known crack in that framing: content similarity and behavioural similarity often disagree. Topically similar items can serve different audiences, and items frequently consumed together can share little surface content. Content-only quantisation yields codes that make semantic sense yet misalign with the recommendation objective. LETTER, MMGRec and others inject collaborative signal into the pipeline, but collaboration stays auxiliary supervision. This paper from Huawei Ireland Research Centre (CIKM 2026) pushes further: SID construction is recursive clustering at heart, the object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal, and SID assignment becomes hierarchical graph partition.

Method

GrIS (Graph-Informed Semantic IDs) splits construction into two independently configurable axes:

In these coordinates, RQ-VAE (TIGER) and RQ-KMeans become the degenerate case of an empty graph, while LETTER, MMGRec, S2GR and spectral clustering each occupy a cell. Prior work is absorbed, not displaced. Two new instantiations:

The two sit in different families: RecDMoN is a global differentiable soft assignment; RQ-GAE is quantisation-style greedy commitment plus graph regularisation.

Results

Every method shares the same TIGER (T5) generative backbone, so gains are attributable to SID construction alone. Residual-quantisation methods use four codebooks of 256 entries, averaged over three seeds.

MethodDatasetMetricResult
RecDMoN vs LETTERToysH@108.16 vs 5.37 (+52.0%)
RecDMoN vs LETTERBeautyH@107.96 vs 6.18 (+28.9%)
RecDMoN vs LETTERSportsH@104.72 vs 3.42 (+38.0%)
RQ-GAE vs S2GR / LETTERBooksH@104.60, +16.5% / +34.1%

RecDMoN leads all evaluated methods on the three small Amazon benchmarks, and RQ-GAE is second on every metric on Yelp. On Books, RecDMoN could not run at all: the DMoN implementation stores a dense adjacency matrix and exceeded GPU memory, while RQ-GAE handled the scale and posted the strongest result there. On MIND the winner is HSTU, a sequential model with no SIDs; the authors attribute this to MIND's long histories (26.6 steps on average) and compact item set, which favour strong sequential models.

The ablations carry weight. For RQ-GAE, the graph contrastive loss alone improves all three datasets tested; graph-propagated inputs alone are mixed, up on Toys (5.38 vs 5.24) but down on Beauty (5.66 vs 5.84); combined is best. RecDMoN beats spectral clustering on balance rather than cluster count: Gini indices of cluster sizes across levels L1-L4 on Beauty run 0.099 to 0.269 versus 0.736 to 0.845 for spectral, whose few dominant clusters waste the code space. The graph-construction ablation, run only on Beauty with windowed co-occurrence, favours deeper thinner hierarchies for RecDMoN (4 levels, branching factor 6) and stronger neighbourhood mixing for RQ-GAE; graph statistics show no monotonic relation with downstream quality. Swapping sentence-t5-base for Qwen3-Embedding-0.6B lifts both methods, RecDMoN more so (Toys H@10 8.16 vs 7.29); RQ-GAE gains less, suggesting graph signal partly compensates for weaker text embeddings.

Why it matters

Teams building generative recommenders get a composable coordinate system: stronger embeddings, denser graphs, and better partitioners can be upgraded and evaluated independently instead of rebuilding the SID pipeline each time. RQ-GAE is the practical insertion point, grafting directly onto an existing RQ-VAE pipeline. Per-dataset numbers are incremental; the durable contribution is the framework that turns scattered methods into comparable instances, onto which improvements on either axis stack.

Limitations

Stated by the authors: collaborative signal drifts as behaviour evolves, so derived partitions go stale, and incremental updates are left to future work; RecDMoN's dense-adjacency implementation caps it at small catalogues, with no sparse or mini-batch variant yet; random-walk and session-level graph construction are untested.

A close read raises more. The strongest model on MIND uses no SIDs, and the paper never charts where SID-based generative recommendation actually wins. Headline numbers match the Qwen3-Embedding configuration exactly, while the text does not state which embedder the baselines used; with sentence-t5-base restored, RecDMoN still reaches 7.29 vs 5.37 on Toys, so the gain is not purely the embedder, but the check matters. And although the framework unifies RecDMoN and RQ-GAE, it offers no rule for choosing a partitioner in a new setting; graph quality is currently trial and error.

Terms

Source

What people are saying

Related papers

All paper explainers