热点事件持续更新
Universal Attention 实现 10 倍 KV-Cache 压缩
1 篇报道1 个报道来源2 小时前 更新
先了解这件事
AI 综述
2026 年 10 月 8 日,arXiv·Machine Learning Theory 发布一手论文《A Self-Pruning Transformer:Universal Attention 实现极限 KV-Cache 压缩》。论文提出 Universal Attention,一种端到端可训练的 Transformer 架构,采用复合衰减机制作为自适应剪枝准则,在保留 RoPE 位置嵌入与 Softmax attention 的同时,移除对 attention 计算贡献最小的 token,以实现极限 KV-Cache 压缩。此前该事件标题称其实现 10 倍 KV-Cache 压缩,但本篇报道正文未给出具体压缩倍数,仅描述机制与目标。
AI 根据报道生成 · 2 小时前更新
最新进展10月8日 12:00
arXiv 论文公开 Universal Attention 架构,用复合衰减自适应剪枝以压缩 KV-Cache。报道时间线
沿着报道,了解事件的不同侧面。
10月8日
- arXiv · Machine Learning TheoryA Self-Pruning Transformer:Universal Attention 实现极限 KV-Cache 压缩
论文提出 Universal Attention,一种端到端可训练的 Transformer 架构,用复合衰减机制作为自适应剪枝准则,在保留 RoPE 位置嵌入与 Softmax attention 的同时移除对 attention 计算贡献最小的 token。
本事件热度走势
还没有足够的连续观测数据,暂不绘制趋势。