Skip to content
Hot eventLive

T2SPO: Trajectory-to-Step Policy Optimization for Agentic RL

1 reports1 sources5 hr ago updated

Get the story

AI overview

2026-10-02,arXiv 发表 T2SPO,面向智能体强化学习的轨迹到步骤策略优化方法。该方法利用成功交互轨迹提供步骤级反馈,通过 TabPFN 回归器估计剩余距离并生成辅助信用信号,在 ALFWorld 和 WebShop 上使用 1.5B 与 7B 语言模型实验,相比 GRPO 一致提升任务成功率。

Generated from reports · updated 4 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Machine Learning Theory
    T2SPO:面向智能体强化学习的轨迹到步骤策略优化

    T2SPO 利用成功交互轨迹提供步骤级反馈,通过 TabPFN 回归器估计剩余距离并生成辅助信用信号,在 ALFWorld 和 WebShop 上用 1.5B 与 7B 语言模型实验,相比 GRPO 一致提升任务成功率。

Heat trend

Current heat 9·Comparable peak 10(Oct 2)·Comparable change over 24 hours –

02.557.510Oct2Oct2Oct2Oct2

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.