Skip to content
Hot eventLive

PowerBench:评测模型在权力转移请求中的偏差

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,研究者在 arXiv(Machine Learning Theory)发布 PowerBench,用于评测语言模型对权力转移请求的拒绝偏差。该基准区分自我赋能、削弱他人权力和夺取权力三类请求,并设置不转移权力的对照请求。团队开源了覆盖权力领域、情境、受影响方规模和用户既有权力地位的数据集,评测 24 个模型(12 个来自美国开发者、12 个来自中国开发者),并在用户与受影响方国籍互换、AI 智能体作为用户、8 种请求语言三种条件下进行测试。目前报道仅涉及基准发布与评测设置,未披露具体评测结果。

Generated from reports · updated 15 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Machine Learning Theory
    PowerBench:评测语言模型在权力转移请求中的偏差

    研究者发布 PowerBench,用于评测语言模型对权力转移请求的拒绝偏差,区分自我赋能、削弱他人权力和夺取权力三类,并设置不转移权力的对照请求。团队开源了覆盖权力领域、情境、受影响方规模和用户既有权力地位的数据集,评测 24 个模型(12 个来自美国、12 个来自中国开发者),在用户与受影响方国籍互换、AI 智能体作为用户、8 种请求语言三种条件下测试。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.