Skip to content
Hot eventLive

DaCe-DT:面向异构任务的离线多任务强化学习框架

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv Machine Learning Theory(一手来源)发布论文,提出离线多任务强化学习框架 DaCe-DT,从数据层面解决三个瓶颈:提示长度利用不足、随机采样提示段语义无关、碎片化轨迹误导监督。框架包含长度门控提示掩码(LGPM)、检索增强提示构建(RAPC)和价值自适应回报校准(VARC)三个组件。Meta-World 实验显示,DaCe-DT 在最优数据集上平均提升11.73%,在次优数据集上提升13.34%,优于现有最先进方法。目前报道仅此一篇,无更早或相互矛盾的信息。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Machine Learning Theory
    DaCe-DT:面向异构任务的离线多任务强化学习框架

    论文提出离线多任务强化学习框架 DaCe-DT,从数据层面解决提示长度利用不足、随机采样提示段语义无关、碎片化轨迹误导监督三个瓶颈。框架包含长度门控提示掩码(LGPM)、检索增强提示构建(RAPC)和价值自适应回报校准(VARC)。Meta-World 实验显示,DaCe-DT 在最优数据集上平均提升 11.73%,在次优数据集上提升 13.34%,优于现有最先进方法。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.