跳到正文
原文
arXiv · Machine Learning Theory· Irving Giovani Bronzatti Petrazzini, Eric Aislan Antonelo·· 3 小时前AI 评分22

Disagreement-Regularized Imitation Learning(DRIL)用于基于图像的连续控制:高斯与 Beta 策略研究

Disagreement-Regularized Imitation Learning for Image-Based Continuous Control with Gaussian and Beta Policies

AI 导读

DRIL 将克隆策略间的分歧转化为强化学习奖励,在基于图像的连续控制任务中提升少样本表现。CarRacing 实验显示,score-selected DRIL 在仅 1 条轨迹时分别相对最优行为克隆提升 61%(clipped-action)和 112%(bounded-action);20 条轨迹时优势收窄,Bounded Beta 行为克隆仍高出约 7%。

来源:arXiv · Machine Learning Theory · arxiv.org