Skip to content
Hot eventLive

FERPO:前向熵正则化策略优化

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026-10-02,arXiv 发表 FERPO,一种在线最大熵强化学习算法。该方法通过 critic 值进行策略改进,无需对 critic 求动作梯度;利用熵与 KL 散度正则化推导目标动作分布,并借助前向 KL 目标与自归一化重要性采样(SNIS)拟合策略,鼓励覆盖多个高价值模式以促进探索。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Statistics Machine Learning
    FERPO:前向熵正则化策略优化

    FERPO 是一种在线最大熵强化学习算法,通过 critic 值进行策略改进,无需对 critic 求动作梯度。算法利用熵和 KL 散度正则化推导目标动作分布,并通过前向 KL 目标与自归一化重要性采样(SNIS)拟合策略,鼓励覆盖多个高价值模式以促进探索。

Heat trend

There is not enough continuous observation data to draw a trend yet.