Hot eventLive
SWAP:VLA模型逐步动作策略路由框架
1 reports1 sources3 hr ago updated
Get the story
AI overview
研究者提出 SWAP(StepWise Action Policy Routing)框架,面向视觉-语言-动作(VLA)模型,在执行过程中动态组合多个 VLA 策略。该框架将策略路由建模为离线强化学习问题,由路由评论器根据当前观测在每一步选择最合适的策略。目前报道仅涉及该框架的提出与基本机制,未见后续实验结果或应用进展。
Generated from reports · updated 3 hr ago
LatestOct 7
研究者提出SWAP框架,以离线强化学习路由评论器在每步动态选择并组合多个VLA策略。Timeline
Follow the coverage from different angles.
Oct 7, 2026
- arXiv · RoboticsSWAP:面向视觉-语言-动作模型的逐步动作策略路由
研究者提出 SWAP(StepWise Action Policy Routing)框架,在执行中动态组合多个 VLA 策略,将策略路由建模为离线强化学习问题,由路由评论器根据当前观测在每步选择最合适的策略。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.