Scaling Diverse Language Generation for 3D Visual Grounding

arXiv:2606.20946v1 Announce Type: new Abstract: Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for enabling agents to correspond spatial language with objects in the physical world. However, the lack of diverse descriptions at scale prevents models from generalizing beyond simple linguistic patterns. Recent such attempts lack diversity in the constraint types and language used to ground objects. Capt...

arXiv cs.CL ·Austin T. Wang, Dongchen Yang, Angel X. Chang ·
compartilhar: