Skip to content
arXiv · Computation and Language· Yang Li, Gongle Xue, Yuheng Yuan, Yijia Guo, Shizhe Zhang, Liwen Hu, Lei Ma·· 7 hr agoSelectedAI score60

诊断推理语言模型的 On-Policy Self-Distillation

Diagnosing On-Policy Self-Distillation for Reasoning Language Models

AI brief

论文诊断了数学推理场景下的 On-Policy Self-Distillation(OPSD),覆盖 0.6B–8B 参数模型。实验与 token 级分析表明,教师信号受推理模式对齐与完整教师前缀影响,OPSD 仅在狭窄兼容 regime 中有效,否则出现长度膨胀、稳定退化或行为崩溃,且教师信号不稳定、无法预测下游性能。

Why it matters

论文通过受控实验和 token 级分析指出 OPSD 并非通用可靠的推理改进后训练方法。

Source: arXiv · Computation and Language · arxiv.org