Skip to content
Hot eventLive

法律思维链忠实性反事实审计研究发布

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv Computers and Society 发布一项反事实审计研究(一手报道)。论文针对七款开源大语言模型(参数规模8B-70B),在四个法律推理基准(CaseHOLD、ECHR、SCOTUS、ContractNLI)上开展实验:固定案情后替换模型生成中引用的法律依据。结果显示,模型在66.7%-100%的生成中能正确命名所引用的法律依据,但判决结果随依据改变的比例较低:CaseHOLD 为0.0%-21.7%,ECHR 与 SCOTUS 为30.0%-76.7%,ContractNLI 为43.3%-50.0%。这表明模型引用的法律依据常与判决结果脱钩,即思维链中引用的依据对最终结论的忠实性有限。研究覆盖的基准与模型范围、正确命名依据的比例区间以及各基准判决随依据改变的比例区间,均以报道披露为准。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Computers and Society
    反事实审计显示大语言模型引用的法律依据常与判决结果脱钩

    论文对七款开源模型(8B-70B)在四个法律推理基准上做反事实审计,固定案情后替换被引用的法律依据,发现模型在66.7%-100%的生成中能正确命名依据,但判决随依据改变的比例仅0.0%-21.7%(CaseHOLD)、30.0%-76.7%(ECHR与SCOTUS)和43.3%-50.0%(ContractNLI)。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.