长上下文混合模型机理研究:从混合注意力到混合位置
Get the story
2026年10月8日,arXiv Computation and Language 发表一手论文《长上下文混合模型机理 1.1:从混合注意力到混合位置》,提出长上下文混合模型机理系列。论文比较全注意力与滑动窗口注意力(SWA)、线性注意力(LA,含 GLA 与 GDN)的混合架构,发现上下文扩展中的跷跷板效应:LA 混合更受益于长上下文持续预训练,SWA 混合在长度外推上更好;SWA 存在短上下文学习陷阱、短窗倦怠与长窗懒惰。论文提出滑动窗口线性注意力,实现 16 倍免训练长度外推,在 64k 上下文 NIAH-SK1 保持 100% 准确率。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and Language长上下文混合模型机理 1.1:从混合注意力到混合位置
论文提出长上下文混合模型机理系列,比较全注意力与滑动窗口注意力(SWA)、线性注意力(LA,含 GLA 与 GDN)混合架构。发现上下文扩展中的跷跷板效应:LA 混合更受益于长上下文持续预训练,SWA 混合在长度外推更好;SWA 存在短上下文学习陷阱、短窗倦怠与长窗懒惰。提出滑动窗口线性注意力,实现 16 倍免训练长度外推,在 64k 上下文 NIAH-SK1 保持 100% 准确率。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.