Hot eventLive
T2SPO: Trajectory-to-Step Policy Optimization for Agentic RL
1 reports1 sources5 hr ago updated
Get the story
AI overview
2026-10-02,arXiv 发表 T2SPO,面向智能体强化学习的轨迹到步骤策略优化方法。该方法利用成功交互轨迹提供步骤级反馈,通过 TabPFN 回归器估计剩余距离并生成辅助信用信号,在 ALFWorld 和 WebShop 上使用 1.5B 与 7B 语言模型实验,相比 GRPO 一致提升任务成功率。
Generated from reports · updated 4 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Machine Learning TheoryT2SPO:面向智能体强化学习的轨迹到步骤策略优化
T2SPO 利用成功交互轨迹提供步骤级反馈,通过 TabPFN 回归器估计剩余距离并生成辅助信用信号,在 ALFWorld 和 WebShop 上用 1.5B 与 7B 语言模型实验,相比 GRPO 一致提升任务成功率。
Heat trend
Current heat 9·Comparable peak 10(Oct 2)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.