Airbnb Trains RL Query Suggestion Model Using Search Reranker Scores as Reward
_reachsumit · x · 2026-10-06
Airbnb researchers present SCOUT, a framework for supply-aware cold-start proactive query suggestion in travel search.
Challenges addressed:
- Travel search is constrained by physical inventory: "romantic beachfront villa" yields abundant results in Bali but few in Tokyo, so aligning with user preferences alone isn't grounded in available supply.
- Travel platforms traditionally use faceted search with no free-text queries, creating a cold-start problem: no historical query logs for demand-side alignment and no seed query at request time.
Approach: SCOUT substitutes missing demand-side user feedback with supply-side system feedback—treating the search engine as a reinforcement learning environment and using the search engine's reranker scores as reward to train a generative query suggestion model, so suggestions match available listings without extra calls at serve time.
A notable industrial case of applying LLM-based generative query suggestion to inventory-constrained search.
More from Research
- Survey: 65% of Japanese seniors prefer robot-assisted nursing homes, willing to pay 8% more — HealthcareLdr · 2026-10-06
- AI Scholar Yi Ma Proposes Rolling 10-Paper Cap on arXiv to Curb Paper Flooding — YiMaTweets · 2026-10-06
- FlashDexRetarget: one RL policy retargets hand demos at 90% success, ~100x less compute — KyleMorgenstein · 2026-10-06
- Stanford ACE team unveils Sentry: failure tips in context hurt LLM agents, +39% gains — StanfordAILab · 2026-10-06
- From Kaggle Champion to ULMFiT: How Jeremy Howard Rewrote Language Model Training — bigaiguy · 2026-10-06
- Training on a Post-Trained Model Often 'Fries' It, Causing Reality Drift — Sauers_ · 2026-10-06