Nautil:LLM调查员结案决策研究
Get the story
2026年10月5日,arXiv Computation and Language发布一手论文《LLM调查员何时结案:基于证据的关闭决策研究》。论文研究LLM调查员在事故、缺陷和宕机调查中何时应结案,提出结案准确性、证据依赖性和结论与缺口质量三项评测,并构建含731份审计案例的Nautil数据集。实验显示,微调9B模型后,过度断言从97%降至35%,正确且不夸大的结论从3%升至43%;仅奖励结案决策的强化学习将平衡准确率从69.2提升到83.3。目前未见后续报道或矛盾信息。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and LanguageLLM 调查员何时结案:基于证据的关闭决策研究
论文研究 LLM 调查员在事故、缺陷和宕机调查中何时应结案,提出结案准确性、证据依赖性和结论与缺口质量三项评测,并构建含 731 份审计案例的 Nautil 数据集。微调 9B 模型后,过度断言从 97% 降至 35%,正确且不夸大的结论从 3% 升至 43%;仅奖励结案决策的强化学习将平衡准确率从 69.2 提升到 83.3。
Heat trend
Current heat 6·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.