VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers
arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily on how they are orchestrated: which tools are used, in what order, with what parameters, and under what visual conditions. Existing visual-programming agents typically generate a fixed solution pipeline, making them brittle under dense objects, occlusi...
arXiv cs.CV
·Jinchao Ge, Lingqiao Liu, Shuwen Zhao, Lei Wang
·
// relacionados
Leia também
Blog
Do Medical Foundation Models Generalize on the African Brain?
Blog
A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging
Blog
RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning
Blog