Skip to content
Hot eventLive

XiangqiBench:中国象棋闭环评估LLM Agent基准

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv Computation and Language 发布一手报道,介绍 XiangqiBench——一个用于评估 LLM 智能体的可执行基准。该基准从 119 个有引擎或仅搜索验证的强制杀局出发,要求 LLM 智能体对引擎防守方完成将杀,并用交互式 REPL 区分真实走子、状态查询与前向模拟,以检验智能体在闭环环境下的实际决策能力。目前报道仅涉及该基准的设计与任务设定,未披露具体实验结果或后续更新。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Computation and Language
    XiangqiBench:用中国象棋闭环评估 LLM 智能体

    XiangqiBench 是一个可执行基准,从 119 个有引擎或仅搜索验证的强制杀局出发,要求 LLM 智能体对引擎防守方完成将杀,并用交互式 REPL 区分真实走子、状态查询与前向模拟。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.