World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
arXiv:2607.27599v1 Announce Type: cross Abstract: Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world mode...
arXiv cs.RO
·Xiangcheng Zhang, Yilun Du
·
// relacionados