Skip to content
Hot eventLive

研究:灾难性遗忘集中于新语料未出现的词元输出嵌入

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv Computation and Language 发布一项一手研究,聚焦大语言模型在持续预训练与微调中的灾难性遗忘问题。研究发现,遗忘并非均匀分布,而是集中在新语料中罕见或未被提及的 token 的输出嵌入上;与此同时,模型主体在对应 sqrt(v-hat) 区间保持惰性,变化有限。该结果提示遗忘主要发生在输出嵌入层而非整体参数,为理解持续训练中的知识损失提供了新观察。目前报道仅涉及这一项研究结论,未见后续验证或对比数据。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computation and Language
    数据从未提及的 token 输出嵌入中的灾难性遗忘研究

    一项针对 LLM 持续预训练与微调中灾难性遗忘的研究发现,遗忘集中在新语料中罕见 token 的输出嵌入,而模型主体对应 sqrt(v-hat) 区间保持惰性。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.