跳到正文
原文
Google Developers · AI·· 2 小时前精选AI 评分76

MaxText 复现 Olmo 3 7B 预训练:TPU 大规模训练案例研究

Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs

AI 导读

Google 团队在 Google Cloud TPUs 上用 MaxText 从头复现了 Ai2 的 Olmo 3 7B 预训练与中期退火,覆盖 stage-1(约 5.93T token / 1.41M 步)和 stage-2(47,684 步),在多个 held-out 指标上与 Ai2 参考曲线对齐。

推荐理由

通过独立复现展示了跨框架验证方法与数据加载陷阱,读者可借鉴其 held-out 评估策略和 resume 校验手段。

来源:Google Developers · AI · developers.googleblog.com