Skip to content
Hot eventLive

DART-ES:面向进化策略微调LLM的难度感知重加权与定向回放

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月7日,arXiv Machine Learning Theory 发布一手报道,提出 DART-ES 方法,用难度感知重加权与定向回放改进进化策略(ES)对大语言模型的全参数微调,仅靠前向计算即可训练。报道给出多项实验结果:DART-ES 在五个基座模型上均优于 ES,平均准确率从 72.07% 提升至 73.53%;在 GSM8K 上超过 GRPO 的 73.26%;在五个数学推理基准上平均准确率为 49.20%,高于 ES 的 48.34%。目前报道仅包含上述事实,未见与其他报道矛盾之处,也未见后续进展或独立验证信息。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 7, 2026
  1. arXiv · Machine Learning Theory
    DART-ES:面向进化策略微调 LLM 的难度感知重加权与定向回放

    提出 DART-ES,用难度感知重加权与定向回放改进进化策略(ES)对 LLM 的全参数微调,仅靠前向计算即可训练。DART-ES 在五个基座模型上均优于 ES,平均准确率从 72.07% 提升至 73.53%,超过 GRPO 在 GSM8K 上的 73.26%;在五个数学推理基准上平均准确率为 49.20%,高于 ES 的 48.34%。

Heat trend

Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –

02.557.510Oct7Oct7Oct7Oct7

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.