新论文提出 OPRD:弱教师反向蒸馏让学生模型超越教师上限

algo_diver · x · 2026-09-10

arXiv 论文《Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation》(Aaron Courville 等参与)研究弱到强泛化:新模型代际更替时,能否不从零重训,而是利用旧弱模型加速训练出更强的继任者。

核心方法 On-Policy Reverse Distillation (OPRD):

所属事件:新论文提出 OPRD:在线反向蒸馏实现弱到强泛化(2 条相关)→

原文链接 →

「模型」频道最新

更多「模型」频道 AI 资讯 →