arXiv · Information Retrieval· Sietse Schelpe·· 7 hr agoSelectedAI score79
Galahad 通过字节精确记忆让 LLM 阅读成为一次性成本
Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
AI brief
论文提出 Galahad 内存层,让 vLLM、SGLang 等运行时缓存 KV 状态,避免重复计算。
Why it matters
原文给出具体数字和三种运行时对比,读者可据此评估 Galahad 在能耗与延迟上的实际收益。
Source: arXiv · Information Retrieval · arxiv.org