Skip to content
Hot eventLive

MoE预训练专家耦合减少All-to-All通信

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv Computation and Language 发表一项MoE预训练研究(一手报道),提出相关专家放置与token洗牌两种方法,用于降低all-to-all通信开销;两种方法不改变路由决策和专家参数。研究发现,预训练早期路由器已学会以相关模式分配token,在top-2设置下,0.8%的专家对被42%的token共同选中。目前报道仅披露上述方法与统计结果,未涉及实验基准、训练规模或后续验证进展。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computation and Language
    MoE 预训练中的专家耦合:用相关放置与 token 洗牌降低 all-to-all 开销

    一项 MoE 预训练研究提出相关专家放置与 token 洗牌两种方法,在不改变路由决策和专家参数的前提下降低 all-to-all 通信开销。研究发现预训练早期路由器已学会以相关模式分配 token,top-2 下 0.8% 的专家对被 42% 的 token 共同选中。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.