EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories
arXiv:2607.00020v1 Announce Type: new Abstract: Spatial grounding remains a key limitation of vision-language-action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instructions, they often lack an explicit representation of how objects are arranged in space, including support, containment, ordering, occlusion, and depth-sensitive relations. We introduce EmbodimentSemantic, a spatial scene-graph dataset and benchmark for evaluating relationa...
arXiv cs.RO
·Hassan Jaber, Refinath S N, Luca Cagliero, Christopher E. Mower, Haitham Bou-Ammar
·
// relacionados
Leia também
Editorial
RynnValue: o relógio do vídeo como recompensa para robôs
Blog
Anthropic aplica marca d'água a todas as saídas do Claude globalmente, com marcas que "podem persistir mesmo após alguma edição"
Blog
webAI lança TwIL-LM: uma família de modelos de lógica formal de 1,7B e 3B para autoformalização em hardware local
Blog