Hot eventLive
XiangqiBench:中国象棋闭环评估LLM Agent基准
1 reports1 sources16 hr ago updated
Get the story
AI overview
2026年10月5日,arXiv Computation and Language 发布一手报道,介绍 XiangqiBench——一个用于评估 LLM 智能体的可执行基准。该基准从 119 个有引擎或仅搜索验证的强制杀局出发,要求 LLM 智能体对引擎防守方完成将杀,并用交互式 REPL 区分真实走子、状态查询与前向模拟,以检验智能体在闭环环境下的实际决策能力。目前报道仅涉及该基准的设计与任务设定,未披露具体实验结果或后续更新。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
Oct 5, 2026
- arXiv · Computation and LanguageXiangqiBench:用中国象棋闭环评估 LLM 智能体
XiangqiBench 是一个可执行基准,从 119 个有引擎或仅搜索验证的强制杀局出发,要求 LLM 智能体对引擎防守方完成将杀,并用交互式 REPL 区分真实走子、状态查询与前向模拟。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.