arXiv · Information Retrieval· Xu Yang, Jiapeng Zhang, Yuxin Chen, Feiqiang Sun, Chengguang Xu, Feng Jin, Zhuo Tang·· 3 小时前AI 评分29
Self-Indexing Attention:兼容压缩的稀疏长上下文 LLM 推理
Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference
AI 导读
研究者提出 Self-Indexing Attention,一种免训练框架,基于共享变换域符号-幅度表示,用关键符号作为可复用的 token 级索引,统一 prefill 分组选择与 decode 检索。
来源:arXiv · Information Retrieval · arxiv.org