Hot eventLive
研究:灾难性遗忘集中于新语料未出现的词元输出嵌入
1 reports1 sources4 hr ago updated
Get the story
AI overview
2026年10月8日,arXiv Computation and Language 发布一项一手研究,聚焦大语言模型在持续预训练与微调中的灾难性遗忘问题。研究发现,遗忘并非均匀分布,而是集中在新语料中罕见或未被提及的 token 的输出嵌入上;与此同时,模型主体在对应 sqrt(v-hat) 区间保持惰性,变化有限。该结果提示遗忘主要发生在输出嵌入层而非整体参数,为理解持续训练中的知识损失提供了新观察。目前报道仅涉及这一项研究结论,未见后续验证或对比数据。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Computation and Language数据从未提及的 token 输出嵌入中的灾难性遗忘研究
一项针对 LLM 持续预训练与微调中灾难性遗忘的研究发现,遗忘集中在新语料中罕见 token 的输出嵌入,而模型主体对应 sqrt(v-hat) 区间保持惰性。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.