Skip to content
Hugging Face Blog·· 3 hr agoSelectedAI score76

Olmo-core 3 发布:面向大规模 MoE 的开放可扩展训练框架

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

AI brief

AllenAI 发布 Olmo-core 3,重设计 MoE 训练系统以支持万亿参数规模。在专家池从 8 扩至 128、每 token 仅选 4 个专家的配置下,总参数从 4.6B 增至 47B,训练吞吐下降不到 5%;在 8 块 NVIDIA B300 上,47B MoE 吞吐约 2.7 倍提升。

Why it matters

展示了专家池扩展与吞吐量保持的具体数据,为评估大规模 MoE 训练效率提供可对比的基准。

Source: Hugging Face Blog · huggingface.co