RobotAPO:对抗物理偏好优化用于机器人操作视频生成
Get the story
研究者提出 RobotAPO,一种在连续流匹配去噪空间中运行的对抗式物理偏好优化框架,用于提升机器人操作视频生成的物理一致性。为训练与评估该框架,研究者构建了包含 10,000 条样本的偏好数据集 AgiBot-PhysPref,用于隔离条件匹配中的物理违规。在留出的 AgiBot 条件下,RobotAPO 相对最强受控内部基线将物理一致性提升 6.8% hard score 与 10.0% soft score;在真实机器人回放中,任务成功率相对提升 37.4%。目前公开信息仅包含该框架、方法与上述实验结果,未见后续独立验证或对比更新。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · RoboticsRobotAPO:用对抗式物理偏好优化提升机器人操作视频生成
研究者提出 RobotAPO,一种在连续流匹配去噪空间中运行的对抗式物理偏好优化框架,并构建了 10,000 条样本的偏好数据集 AgiBot-PhysPref,用于隔离条件匹配的物理违规。在留出的 AgiBot 条件下,RobotAPO 相对最强受控内部基线将物理一致性提升 6.8% hard score 与 10.0% soft score,真实机器人回放中任务成功率相对提升 37.4%。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.