Meituan LongCat's Neighborhood OPSD boosts math reasoning by up to 2.75 Average@12 points on Qwen3
meituan-longcat · hf · 2026-10-02
Meituan LongCat introduces Neighborhood OPSD, an improvement over on-policy self-distillation for math reasoning: a pool of locally perturbed frozen experts supplies complementary reference-aligned corrections, with offline greedy selection and online routing (MaxPeak + quantile selection). Across AIME 2024/2025 and HMMT Feb 2025 it improves Average@12 over OPSD by 2.75/1.67/1.94 points on Qwen3-1.7B/4B/8B, with inference using only the distilled student.
More from Research
- AMap open-sources ABot-Recon: streaming 3D reconstruction from video with a 12-frame local context — rsasaki0109 · 2026-10-02
- Neuralink pretrains decoders on 50,000+ hours of neural data—dataset may be the real moat — CurieuxExplorer · 2026-10-02
- CyberGym cybersecurity benchmark effectively saturated on verified task subset — aryaman2020 · 2026-10-02
- SemEval-2027 Task 9 calls for teams on 4-language multimodal news framing analysis — preslav_nakov · 2026-10-02
- Alternating prompt and model upgrades lift science agent from 42% to 73% — rohanpaul_ai · 2026-10-02
- ISMIR paper teaches a transformer to play in 12 jazz piano legends' styles — umpedronosapato · 2026-10-02