RPTune: Learned Context Curation for LLM Catalog Search
Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan Ö. Arık
cs.IR, cs.CL, cs.LG
2026-10-01
A learned catalog curator (prune + reorder) plus GRPO post-training lifts in-context LLM search on SMB catalogs: +13.5 EM avg across 9 backbones; gemma-4-E4B-it 10.7% to 31.0%.
Product search as currently built serves marketplaces with millions of items, dedicated ML teams, and abundant behavioral logs. Small merchant businesses (SMBs) have none of that. Among 56 real Shopify storefronts analyzed in the paper, 92.9% of catalogs fit within 1M tokens and the median is just 51K, so a long-context LLM can hold the entire inventory in its prompt and pick a product directly, skipping retrieval. With gemini-3.7-flash as the shared backbone, full-catalog prompting beats every retrieval baseline by at least 9.0 EM points on average at 8 seconds latency.
Fitting the catalog is not the same as using it. LLMs use long context unevenly, and relevant items buried mid-context get used poorly (the lost-in-the-middle effect). The paper splits the problem in two: how to curate and lay out the catalog, and how to adapt the LLM to curated contexts.
RPTune (Rank, Prune, Finetune) has three trained pieces, all supervised by synthetic data generated from the catalog alone. No merchant labels, no user logs.
Evaluation covers 7 real merchants (catalogs of 37.5K to 127K tokens) with 100 complex conversational queries each, 10 runs per method.
| Setting | Metric | Result |
| Full-catalog vs best retrieval baseline (gemini-3.7-flash) | EM avg | 36.8 vs 27.8 |
| Curation on 9 frozen backbones | EM / FR | +13.5 / +10.1 avg; all 63 combos improve |
| grok-4.1-fast-non-reasoning + curation | EM | 12.8% to 31.1% |
| gemma-4-E4B-it full pipeline | EM / FR | 10.7% to 31.0%; FR 56.7 to 74.4 |
| Post-training increment | EM | +10.3 avg on top of curation |
| End-to-end latency, 9 LLMs | seconds | -28% avg, up to -71% |
The largest single gain is 31.4 points (grok-4.1-fast-non-reasoning on Beauty Bakerie, 14.8 to 46.2). At the same 25% retention budget, Jina Reranker 3.5 hurts accuracy on 5 of 7 merchants and Gemini Embedding 2 on 3, while RPTune improves all 7, so the gains come from learned curation rather than shorter contexts. The sharpest ablation: replacing the RL reorganizer objective with supervised listwise cross-entropy on relevance labels caps EM at 16.4%, no better than the encoder alone and 4.3 points below RPTune. Generalization holds too: after injecting new products until the catalog doubles, post-trained RPTune loses 14.7% relative EM versus 29.5% for the baseline, and all 7×7 cross-merchant transfers are positive (+18.1 EM average out of domain). Per Figure 1, curation buys accuracy comparable to a model-tier upgrade at up to 21× speedup, and post-training brings the 8B gemma into the accuracy range of frontier models.
This is a worked recipe for replacing a retrieval stack with direct long-context reasoning wherever the candidate set fits in context: small storefronts, internal selection tasks, agents choosing among tens or hundreds of tools or products. The division of labor is worth copying: a 300M encoder plus a small MLP does the context engineering, the LLM reads 25% of the catalog, and latency drops instead of rising. Training needs only catalog metadata, zero human labels. The context-relative reward is reusable in any pipeline that prunes context before RL.
The honest framing: contrastive encoders, GRPO, and position sensitivity are all known pieces. The contribution is the composition and the task framing, a solid but incremental engineering argument.
The paper has no explicit limitations section; these are reader-observed.