跳到正文
原文
arXiv · Information Retrieval· Xu Yang, Jiapeng Zhang, Yuxin Chen, Feiqiang Sun, Chengguang Xu, Feng Jin, Zhuo Tang·· 3 小时前AI 评分29

Self-Indexing Attention:兼容压缩的稀疏长上下文 LLM 推理

Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference

AI 导读

研究者提出 Self-Indexing Attention,一种免训练框架,基于共享变换域符号-幅度表示,用关键符号作为可复用的 token 级索引,统一 prefill 分组选择与 decode 检索。

来源:arXiv · Information Retrieval · arxiv.org