Airbnb Trains RL Query Suggestion Model Using Search Reranker Scores as Reward

_reachsumit · x · 2026-10-06

Airbnb researchers present SCOUT, a framework for supply-aware cold-start proactive query suggestion in travel search.

Challenges addressed:

Approach: SCOUT substitutes missing demand-side user feedback with supply-side system feedback—treating the search engine as a reinforcement learning environment and using the search engine's reranker scores as reward to train a generative query suggestion model, so suggestions match available listings without extra calls at serve time.

A notable industrial case of applying LLM-based generative query suggestion to inventory-constrained search.

Original post →

More from Research

Research channel →