DART-ES:面向进化策略微调LLM的难度感知重加权与定向回放
Get the story
2026年10月7日,arXiv Machine Learning Theory 发布一手报道,提出 DART-ES 方法,用难度感知重加权与定向回放改进进化策略(ES)对大语言模型的全参数微调,仅靠前向计算即可训练。报道给出多项实验结果:DART-ES 在五个基座模型上均优于 ES,平均准确率从 72.07% 提升至 73.53%;在 GSM8K 上超过 GRPO 的 73.26%;在五个数学推理基准上平均准确率为 49.20%,高于 ES 的 48.34%。目前报道仅包含上述事实,未见与其他报道矛盾之处,也未见后续进展或独立验证信息。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Machine Learning TheoryDART-ES:面向进化策略微调 LLM 的难度感知重加权与定向回放
提出 DART-ES,用难度感知重加权与定向回放改进进化策略(ES)对 LLM 的全参数微调,仅靠前向计算即可训练。DART-ES 在五个基座模型上均优于 ES,平均准确率从 72.07% 提升至 73.53%,超过 GRPO 在 GSM8K 上的 73.26%;在五个数学推理基准上平均准确率为 49.20%,高于 ES 的 48.34%。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.