Skip to content
arXiv · Information Retrieval· Sietse Schelpe·· 8 hr agoSelectedAI score79

Galahad 通过字节精确记忆让 LLM 阅读成为一次性成本

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

AI brief

论文提出 Galahad 内存层,让 vLLM、SGLang 等运行时缓存 KV 状态,避免重复计算。

Why it matters

原文给出具体数字和三种运行时对比,读者可据此评估 Galahad 在能耗与延迟上的实际收益。

Source: arXiv · Information Retrieval · arxiv.org