Xiaohongshu Multimodal Inference Optimization: Vision Token Compression & MoE Re-routing

小红书技术REDtech · wechat · 2026-08-05

Xiaohongshu's tech team detailed their dual-end optimization for multimodal LLM inference in the 'Wen Yi Wen' feature, tackling vision token bloat and MoE decode overhead.

Dynamic Vision Token Compression (VisionZip)

Similarity-based Expert Re-routing (SERE)

Both methods require no weight updates and are integrated into vLLM with custom Triton/CUDA kernels, fully compatible with PD disaggregation.

Original post →

More from Infra

Infra channel →