EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories
arXiv:2607.00020v1 Announce Type: new Abstract: Spatial grounding remains a key limitation of vision-language-action (VLA) systems for robotic manipulation. While current models can recognize objects and follow language instructions, they often lack an explicit representation of how objects are arranged in space, including support, containment, ordering, occlusion, and depth-sensitive relations. We introduce EmbodimentSemantic, a spatial scene-graph dataset and benchmark for evaluating relationa...
arXiv cs.RO
·Hassan Jaber, Refinath S N, Luca Cagliero, Christopher E. Mower, Haitham Bou-Ammar
·
// relacionados
Leia também
Blog
Orquestração agêntica: as organizações de IA corporativa têm um problema de implantação, não um problema de plataforma — e a maioria está chamando chatbots de agentes
Blog
Conheça o GPT-Red: um super-hacker LLM que a OpenAI criou para tornar seus modelos mais seguros
Blog
NVIDIA e Japão levam IA e robótica full-stack para todos os setores
Blog