Xiaohongshu & NVIDIA build GR-Inference engine, doubling throughput for Beam Search

小红书技术REDtech · wechat · 2026-09-01

Xiaohongshu's search uses Generative Retrieval (GR) based on SemanticID, characterized by long context, short decode, and large beam width, which general frameworks struggle to support efficiently. Xiaohongshu and NVIDIA jointly built GR-Inference, a specialized engine for this workload.

Core Design

Core Kernel Optimizations

Performance

Under Qwen3-0.6B, Fixed Beam 900, and E2E Latency <100ms:

Original post →

More from Infra

Infra channel →