Skip to content
Hot eventLive

自参照社会偏好:无需观察他人奖励的多智能体合作

1 reports1 sources3 hr ago updated

Get the story

AI overview

研究者提出自参照社会偏好方法,让每个智能体学习自身奖励模型,并将其应用于其他智能体的观测转移,从而在无需观察他人奖励的情况下评估其结果。该方法在 Escape Room、Clean Up 和 Commons Harvest 三个序列社会困境中验证,智能体在独立学习者无法合作的环境中仍能学会合作行为,并常比拥有真实奖励的智能体更公平地分配联合收益。目前未见与早先报道矛盾之处。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 7, 2026
  1. arXiv · Multiagent Systems
    自参照社会偏好:在不观察他人奖励的情况下实现合作

    研究者提出自参照社会偏好方法,让每个智能体学习自身奖励模型并将其应用于其他智能体的观测转移,从而在无需观察他人奖励的情况下评估其结果。该方法在 Escape Room、Clean Up 和 Commons Harvest 三个序列社会困境中验证,智能体在独立学习者无法合作的环境中仍能学会合作行为,并常比拥有真实奖励的智能体更公平地分配联合收益。

Heat trend

Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –

02.557.510Oct7Oct7Oct7Oct7

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.