Skip to content
Hot eventLive

RLVR导致解法模式坍缩,提出ModeBench与Re:Max

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv Machine Learning Theory(一手)发表论文,研究RLVR后训练中的解法模式坍缩问题。论文提出ModeBench基准,可同时返回正确性与所发现的解法模式,用于测量RLVR后训练中的解法多样性。结果显示,RLVR后训练会把概率集中到更少的正确模式上,即使准确率保持或提升,前沿模型也已高度集中。为此论文提出Re:Max方法,在回放缓冲区中为每个已发现模式存储一个已验证样本并均匀训练,在三个模型规模、两种RL目标与更难任务构造上同时提升策略成功率与成功方式数量。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Machine Learning Theory
    RLVR 中解法模式坍缩的测量与缓解

    论文提出 ModeBench,一个同时返回正确性与所发现解法模式的基准,用于测量 RLVR 后训练中的解法多样性。结果显示 RLVR 后训练会把概率集中到更少的正确模式上,即使准确率保持或提升,前沿模型也已高度集中;为此提出 Re:Max,在回放缓冲区中为每个已发现模式存储一个已验证样本并均匀训练,在三个模型规模、两种 RL 目标与更难任务构造上同时提升策略成功率与成功方式数量。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.