Xiaohongshu & NVIDIA build GR-Inference engine, doubling throughput for Beam Search
小红书技术REDtech · wechat · 2026-09-01
Xiaohongshu's search uses Generative Retrieval (GR) based on SemanticID, characterized by long context, short decode, and large beam width, which general frameworks struggle to support efficiently. Xiaohongshu and NVIDIA jointly built GR-Inference, a specialized engine for this workload.
Core Design
- Request-Centric State: Manages resources at the request level, separating logical state trees from physical storage to avoid fragmentation.
- KV Abstraction: Splits KV Cache into request-level shared ContextKV, Beam-private BeamKV, and BeamPath for topology, resolving the share/isolation conflict.
- Scheduling & Policy: Supports Continuous Batching and three Beam width policies: Fixed, Scheduled, and Score-margin.
- Backend: Uses Bucket-based CUDA Graph to reduce overhead and Meta-data driven Replay for cross-request reuse.
Core Kernel Optimizations
- GR-Decode-Attention: A three-kernel pipeline. K1 uses Tensor Cores for dense ContextKV; K2 uses CUDA Cores for sparse BeamKV; K3 performs Log-sum-exp reduction, balancing throughput and latency.
- GPU SID Trie & Constrained Top-K: Keeps Trie resident on GPU and fuses prefix traversal, Token Gather, and local Top-K into a single kernel, eliminating Host-Device sync overhead.
Performance
Under Qwen3-0.6B, Fixed Beam 900, and E2E Latency <100ms:
- GR-Inference achieves 1.5x2.5x higher peak throughput compared to three SGLang community implementations.
- In Item-constrained scenarios, it improves performance by over 3.2x compared to SGLang.
More from Infra
- DIT launches AI token exchange to route requests, claiming 30–70% cost savings — Div_pradeep · 2026-09-01
- Anthropic commits $80B to cloud capacity in a single month — kimmonismus · 2026-09-01
- ESP32 voice assistant: 8x wake-word model compression with AIMET — carrycooldude · 2026-09-01
- LITE plans VCSEL products for AI interconnects, delayed by 1-2 years — zephyr_z9 · 2026-09-01
- antirez shows DeepSeek v4 Flash vision running fast locally on an M5 Max; Metal/CUDA/ROCm support nearly ready — antirez · 2026-09-01
- Nvidia Earnings: Avoiding Consolidation and Dollars per Gigawatt — Stratechery · 2026-09-01