Skip to content
Hot eventLive

PARCEL:法律幻觉检测声明级基准

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,研究者在arXiv(Computation and Language)发布PARCEL基准,用于检验法律声明是否得到所引用判例支持。数据集基于纽约州上诉法院判决构建,含3,396条括号式声明,标注为Supported、Refuted或Not Found三类。研究以零样本方式评测多个SOTA LLM,最强模型准确率达0.97;但在提供完整判决书时仍会把无支持声明误判为支持。结果显示,缺失支持比直接矛盾更难识别,伪造但看似合理的引用导致性能下降最大。该基准为法律领域幻觉检测提供了声明级评测框架。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Computation and Language
    PARCEL:面向法律模型幻觉检测的三分类声明级评测基准

    研究者提出 PARCEL 基准,用于检验法律声明是否得到所引用判例支持。数据集基于纽约州上诉法院判决构建,含 3,396 条括号式声明,标注为 Supported、Refuted 或 Not Found,并以零样本方式评测多个 SOTA LLM。最强模型准确率达 0.97,但仍在提供完整判决书时把无支持声明误判为支持;缺失支持比直接矛盾更难识别,伪造但看似合理的引用导致性能下降最大。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.