CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation
arXiv:2606.28397v1 Announce Type: new Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy. However, most existing VLN methods relying on large models still adopt an open-loop decision-execution approach, where candidate actions are generated from instructions and observations but are rarely verified or corrected ...
arXiv cs.CV
·Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
·
// relacionados
Leia também
Blog
Orquestração agêntica: as organizações de IA corporativa têm um problema de implantação, não um problema de plataforma — e a maioria está chamando chatbots de agentes
Blog
Soofi Consortium lança o Soofi S 30B-A3B: um modelo de fundação MoE híbrido Mamba-Transformer aberto para alemão e inglês
Modelo
thinkingmachines/Inkling
Blog