Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning

arXiv:2606.30686v1 Announce Type: new Abstract: Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot manipulation benchmarks. These gains are commonly interpreted as evidence that semantic representations learned from internet-scale data transfer to physical execution generalization. This position paper argues that the assumption underlying this interpretation -- that semantic generalization is sufficient to su...

arXiv cs.RO ·Taozhao Chen, Ian Manchester, Huaming Chen ·
compartilhar: