SR-OPD:用成功轨迹参照选择蒸馏样本
Get the story
研究者提出 Success-Referenced On-Policy Distillation(SR-OPD),用于在策略蒸馏场景下选择教师监督。该方法按提示与 rollout 是否有成功样本来分配教师 token,优先处理与成功参照隐藏状态轨迹持续分歧的失败 rollout。2026-10-05 该工作以“SR-OPD:按成功参照分配教师 token 的在策略蒸馏”为题发表于 arXiv(Artificial Intelligence 分类,属一手来源)。目前公开信息仅涉及方法思路,尚无实验数据、代码或第三方复现报道。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Artificial IntelligenceSR-OPD:按成功参照分配教师 token 的在策略蒸馏
研究者提出 Success-Referenced On-Policy Distillation(SR-OPD),在在策略蒸馏中按提示与 rollout 是否有成功样本选择教师监督,优先处理与成功参照隐藏状态轨迹持续分歧的失败 rollout。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.