Skip to content
Hot eventLive

LLM时间表达抽取多维泛化能力评测

1 reports1 sources16 hr ago updated

Get the story

AI overview

该研究系统评测多个模型族、架构与推理策略在时间与事件表达抽取任务中的四维泛化表现。结果显示,基础任务表现强通常预示更好泛化,但在显著分布偏移下该关系减弱。归纳式 prompting 在领域偏移、对抗扰动、组合性和长度增加上表现最稳定;规模、架构及演绎与溯因 prompting 的收益不均衡且依赖维度,LLM 泛化无法由单一维度可靠预测。论文被 AACL-IJCNLP 2026 Findings 接收。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Computation and Language
    大语言模型时间抽取任务多维泛化能力评测

    该研究系统评测多个模型族、架构与推理策略在时间与事件表达抽取任务中的四维泛化表现,发现基础任务表现强通常预示更好泛化,但在显著分布偏移下该关系减弱。归纳式 prompting 在领域偏移、对抗扰动、组合性和长度增加上表现最稳定,规模、架构及演绎与溯因 prompting 的收益不均衡且依赖维度,LLM 泛化无法由单一维度可靠预测。论文被 AACL-IJCNLP 2026 Findings 接收。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.