提出MASA:稀疏注意力应建模为矩阵近似
Get the story
2026年10月9日,arXiv Computation and Language 发表论文,提出 Matrix Approximation Sparse Attention(MASA)。论文主张稀疏注意力应建模为矩阵近似,而非仅保留注意力矩阵中的大标量或高质量区域。MASA 采用闭式评分,衡量每个稀疏单元对矩阵乘积近似误差的减少量,可作为插件式修正加入现有稀疏注意力框架,无需改动稀疏核或预算。实验在多种稀疏注意力方法、基准和 LLM backbone 上取得一致精度提升。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and Language稀疏注意力是矩阵近似,而非从数值袋中挑选
论文提出 Matrix Approximation Sparse Attention(MASA),主张稀疏注意力应建模为矩阵近似,而非仅保留注意力矩阵中的大标量或高质量区域。MASA 用闭式评分衡量每个稀疏单元对矩阵乘积近似误差的减少量,可作为插件式修正加入现有稀疏注意力框架,无需改动稀疏核或预算。实验在多种稀疏注意力方法、基准和 LLM backbone 上取得一致精度提升。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.