Walmart applies DPO to Doc2Query, cutting 50% of irrelevant search queries
_reachsumit · x · 2026-10-06
A Walmart team published an arXiv paper (QGDPO) applying direct preference optimization (DPO) to Doc2Query document expansion for e-commerce search.
- Doc2Query often hallucinates irrelevant queries or repeats document content; QGDPO fine-tunes a seq2seq base model, scores predictions with a relevance model, and builds win/lose preference pairs for DPO training.
- The pipeline also uses the relevance model to filter poor predictions before indexing.
- Results: 50% of irrelevant predictions eliminated vs Doc2Query baselines, with a relevance filter removing an additional 14.61%.
- Deployed at full production traffic on walmart.com, with substantial gains in relevance and user engagement.
More from Research
- AI-Discovered Algorithm Delivers First True Subquadratic 3SUM and Subcubic APSP — FrnkNlsn · 2026-10-06
- Amazon AGI's ALoDLM Beats Diffusion and AR Baselines, Hits 612 tok/s at 8B — arankomatsuzaki · 2026-10-06
- Schmidhuber: His 1991 Paper Introduced Pre-training, Positional Encoding and Distillation — SchmidhuberAI · 2026-10-06
- Neuroscientist argues experimental neuroscience can't crack intelligence, but aids disease therapies — aran_nayebi · 2026-10-06
- Scientists unveil AI that can recreate exactly what you're looking at — ChuckDBrooks · 2026-10-06
- FinePhrase (COLM Oral): 1T-Token Study Finds Structured Synthetic Data Beats Curated Web, Cuts Costs 30x — edwardbeeching · 2026-10-06