Universal Attention 实现 10 倍 KV-Cache 压缩
Get the story
2026 年 10 月 8 日,arXiv·Machine Learning Theory 发布一手论文《A Self-Pruning Transformer:Universal Attention 实现极限 KV-Cache 压缩》。论文提出 Universal Attention,一种端到端可训练的 Transformer 架构,采用复合衰减机制作为自适应剪枝准则,在保留 RoPE 位置嵌入与 Softmax attention 的同时,移除对 attention 计算贡献最小的 token,以实现极限 KV-Cache 压缩。此前该事件标题称其实现 10 倍 KV-Cache 压缩,但本篇报道正文未给出具体压缩倍数,仅描述机制与目标。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Machine Learning TheoryA Self-Pruning Transformer:Universal Attention 实现极限 KV-Cache 压缩
论文提出 Universal Attention,一种端到端可训练的 Transformer 架构,用复合衰减机制作为自适应剪枝准则,在保留 RoPE 位置嵌入与 Softmax attention 的同时移除对 attention 计算贡献最小的 token。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.