跳到正文
原文
arXiv · Information Retrieval· Sietse Schelpe·· 2 小时前精选AI 评分79

Galahad 通过字节精确记忆让 LLM 阅读成为一次性成本

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

AI 导读

论文提出 Galahad 内存层,让 vLLM、SGLang 等运行时缓存 KV 状态,避免重复计算。

推荐理由

原文给出具体数字和三种运行时对比,读者可据此评估 Galahad 在能耗与延迟上的实际收益。

来源:arXiv · Information Retrieval · arxiv.org