arXiv · Robotics· Masatoshi Tateno, Takehiko Ohkawa, Yueh-Hua Wu, Hanlong Li, Tatsuya Matsushima, Yoichi Sato, Kei Ota·· 4 hr agoAI score34
YUBI-STAG:通过自动化视频-语言定位实现 VLA 的接触与语义丰富对齐
YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding
AI brief
YUBI-STAG 框架通过接触目标分割结合视觉语言模型,自动为机器人操作演示补充物体身份、属性、状态、单手夹爪动作、双臂协调及空间定位交互等细粒度语义,并将预训练 VLA 与精细操作语言对齐。
Source: arXiv · Robotics · arxiv.org