Meituan LongCat's Neighborhood OPSD boosts math reasoning by up to 2.75 Average@12 points on Qwen3

meituan-longcat · hf · 2026-10-02

Meituan LongCat introduces Neighborhood OPSD, an improvement over on-policy self-distillation for math reasoning: a pool of locally perturbed frozen experts supplies complementary reference-aligned corrections, with offline greedy selection and online routing (MaxPeak + quantile selection). Across AIME 2024/2025 and HMMT Feb 2025 it improves Average@12 over OPSD by 2.75/1.67/1.94 points on Qwen3-1.7B/4B/8B, with inference using only the distilled student.

Original post →

More from Research

Research channel →