arXiv · Information Retrieval· Xu Yang, Jiapeng Zhang, Yuxin Chen, Feiqiang Sun, Chengguang Xu, Feng Jin, Zhuo Tang·· 4 hr agoAI score29
Self-Indexing Attention:兼容压缩的稀疏长上下文 LLM 推理
Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference
AI brief
研究者提出 Self-Indexing Attention,一种免训练框架,基于共享变换域符号-幅度表示,用关键符号作为可复用的 token 级索引,统一 prefill 分组选择与 decode 检索。
Source: arXiv · Information Retrieval · arxiv.org