Hot eventLive
大规模推荐系统有效训练时间优化研究
1 reports1 sources7 hr ago updated
Get the story
AI overview
大规模推荐训练中生命周期开销导致仅 50-60% 的端到端时间用于新数据训练。作者提出覆盖全训练栈的优化方案,包括通信消除、管道重叠、动态形状处理、自动调优剪枝、可复用 PyTorch 2 编译缓存、异步检查点、独立模型发布和恢复成本降低。ETT% 在每个 benchmark 上平均提升 15.5%,最大工作负载达 85%,全集群从约 80% 提升至 90% 以上。
Generated from reports · updated 6 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Information Retrieval大规模推荐系统有效训练时间优化
大规模推荐训练中生命周期开销导致仅 50-60% 的端到端时间用于新数据训练。作者提出覆盖全训练栈的优化方案,包括通信消除、管道重叠、动态形状处理、自动调优剪枝、可复用 PyTorch 2 编译缓存、异步检查点、独立模型发布和恢复成本降低。ETT% 在每个 benchmark 上平均提升 15.5%,最大工作负载达 85%,全集群从约 80% 提升至 90% 以上。
Heat trend
Current heat 8·Comparable peak 10(Oct 2)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.