VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

arXiv:2608.16978v1 Announce Type: new Abstract: Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action representation it never saw in pretraining, which throws away much of the reasoning that made the model worth reaching for. We go the other way and keep the VLM frozen. It writes the policy as a short Python control function, with no demonstrations and no fine-tuning. Writing that code once is open-loop, though. Existing closed-loop methods r...

arXiv cs.RO ·Dhia Naouali, Minghan Wu, Claudia Wong, Abhinav Puthran, Omar G. Younis ·
compartilhar: