法律思维链忠实性反事实审计研究发布
Get the story
2026年10月9日,arXiv Computers and Society 发布一项反事实审计研究(一手报道)。论文针对七款开源大语言模型(参数规模8B-70B),在四个法律推理基准(CaseHOLD、ECHR、SCOTUS、ContractNLI)上开展实验:固定案情后替换模型生成中引用的法律依据。结果显示,模型在66.7%-100%的生成中能正确命名所引用的法律依据,但判决结果随依据改变的比例较低:CaseHOLD 为0.0%-21.7%,ECHR 与 SCOTUS 为30.0%-76.7%,ContractNLI 为43.3%-50.0%。这表明模型引用的法律依据常与判决结果脱钩,即思维链中引用的依据对最终结论的忠实性有限。研究覆盖的基准与模型范围、正确命名依据的比例区间以及各基准判决随依据改变的比例区间,均以报道披露为准。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computers and Society反事实审计显示大语言模型引用的法律依据常与判决结果脱钩
论文对七款开源模型(8B-70B)在四个法律推理基准上做反事实审计,固定案情后替换被引用的法律依据,发现模型在66.7%-100%的生成中能正确命名依据,但判决随依据改变的比例仅0.0%-21.7%(CaseHOLD)、30.0%-76.7%(ECHR与SCOTUS)和43.3%-50.0%(ContractNLI)。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.