FluidPD:支持SLO感知的预填充-解码分离LLM服务就地弹性伸缩
Get the story
FluidPD 是一个面向 Prefill-Decode 分离式 LLM 推理的系统,主打 SLO 感知的就地弹性。为应对瞬时负载失衡,它通过 FluidToken 将有限比例的 prefill 计算卸载到有空闲余量的 decode worker;为应对持续性失衡,它通过 FluidRole 在不重载模型、不重启引擎的情况下,就地切换运行中 worker 的 prefill/decode 角色。该机制据称可在满足 SLO 的同时提升资源利用弹性。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Artificial IntelligenceFluidPD:面向 SLO 感知 Prefill-Decode 分离式 LLM 推理的就地弹性机制
FluidPD 是一个 Prefill-Decode 分离式 LLM 推理系统,提供 SLO 感知的就地弹性。它通过 FluidToken 将有限比例的 prefill 计算卸载到有空闲余量的 decode worker 处理瞬时失衡,通过 FluidRole 在不重载模型、不重启引擎的情况下就地切换运行中 worker 的 prefill/decode 角色应对持续失衡。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.