PARCEL:法律幻觉检测声明级基准
Get the story
2026年10月9日,研究者在arXiv(Computation and Language)发布PARCEL基准,用于检验法律声明是否得到所引用判例支持。数据集基于纽约州上诉法院判决构建,含3,396条括号式声明,标注为Supported、Refuted或Not Found三类。研究以零样本方式评测多个SOTA LLM,最强模型准确率达0.97;但在提供完整判决书时仍会把无支持声明误判为支持。结果显示,缺失支持比直接矛盾更难识别,伪造但看似合理的引用导致性能下降最大。该基准为法律领域幻觉检测提供了声明级评测框架。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and LanguagePARCEL:面向法律模型幻觉检测的三分类声明级评测基准
研究者提出 PARCEL 基准,用于检验法律声明是否得到所引用判例支持。数据集基于纽约州上诉法院判决构建,含 3,396 条括号式声明,标注为 Supported、Refuted 或 Not Found,并以零样本方式评测多个 SOTA LLM。最强模型准确率达 0.97,但仍在提供完整判决书时把无支持声明误判为支持;缺失支持比直接矛盾更难识别,伪造但看似合理的引用导致性能下降最大。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.