Skip to content
arXiv · Computation and Language· Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon·· 7 hr agoSelectedAI score60

Recovering Off-Policy Supervision for Speculative Decoding

Recovering Off-Policy Supervision for Speculative Decoding

AI brief

论文提出基于 rollout 的训练框架,通过 Anchor-Label Relabelling(ALR)和 In-Rollout Anchors(IRA)恢复被离策略 token 破坏的监督信号。在固定视觉语言和文本语料上,贪婪接受长度较 DFlash 最高提升 36.5%,单轮即超越最佳擦除策略,三轮后达到目标重生成响应的接受长度。代码已开源。

Why it matters

为固定语料上的推测解码训练提供了恢复完整监督的新方法,在接受长度和计算效率上对现有擦除基线形成优势。

Source: arXiv · Computation and Language · arxiv.org