Skip to content
arXiv · Computation and Language· Radha Gulhane, Quentin Anthony, Beren Millidge·· 4 hr agoAI score35

MoE 预训练中的专家耦合:用相关放置与 token 洗牌降低 all-to-all 开销

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

AI brief

一项 MoE 预训练研究提出相关专家放置与 token 洗牌两种方法,在不改变路由决策和专家参数的前提下降低 all-to-all 通信开销。研究发现预训练早期路由器已学会以相关模式分配 token,top-2 下 0.8% 的专家对被 42% 的 token 共同选中。

Source: arXiv · Computation and Language · arxiv.org