LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models

arXiv:2608.17600v1 Announce Type: new Abstract: Visual cues are increasingly adopted to guide robot learning, but whether Vision-Language-Action (VLA) models can reliably follow authorized cues while disregarding unauthorized ones remains unclear. Existing work covers only a narrow range of cue forms and focuses on final task success, providing only a coarse assessment of cue-following capability. Treating all visual cues as authorized also leaves safety risks of unauthorized following unexplore...

arXiv cs.RO ·Zhengyan Qian, Rui Yan, Alex Jinpeng Wang, Jinhui Tang ·
compartilhar: