FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling
arXiv:2607.29596v1 Announce Type: new Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized solution for embodied AI applications. Conventional VLAs usually concentrate on current digital cognition. While some efforts are made to enhance VLAs' reasoning capabilities by capturing temporal information, encoding the long-context history causes an efficiency-decreasing issue. To reconcile the conf...
arXiv cs.RO
·Li Lin, Wujun Xu, Weiwei Meng, Kaiwen Xia, Kang Hao Cheong, Shuai Wang
·
// relacionados
Leia também
Modelo
ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot
Editorial
MiniMax H3: um modelo só para texto, imagem, vídeo e áudio — por um terço do preço
Modelo
sensenova/SenseNova-U1.5-8B-MoT-Preview
Blog