论文:结构化MoE专家选择用于智能体强化学习
Get the story
2026年10月7日,arXiv Machine Learning Theory 发布一项面向智能体强化学习的 MoE 专家选择结构化研究。研究提出一套分层路由控制框架,在 RL 后训练中显式对齐回合级专家选择与智能体操作,并引入 token 级正则化以保持局部一致性。研究发现,现成 MoE 模型的专家路由在语义相似操作(如 READ、UPDATE)的回合之间重叠更高,而标准 RL 算法会让路由失控。为此,该框架在所有评测基准上将成功率提升超过 10 个百分点,并引入熵门控机制来解决后训练稳定性问题。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Machine Learning Theory面向智能体强化学习的 MoE 专家选择结构化研究
一项研究提出面向智能体任务的分层路由控制框架,在 RL 后训练中显式对齐回合级专家选择与智能体操作,并用 token 级正则化保持局部一致性。研究发现现成 MoE 模型的专家路由在语义相似操作(如 READ、UPDATE)回合间重叠更高,而标准 RL 算法会让路由失控。该框架在所有评测基准上将成功率提升超过 10 个百分点,并引入熵门控机制解决后训练稳定性问题。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.