Compiling and Benchmarking Task-State Horizons for Embodied Agents

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long-horizon evaluation, but primarily characterize difficulty through action-sequence length and subtask complexity, overlooking a distinct challenge: agents must track evolving task-relevant world states induced by both their exploration and environmental dynamics. We define the span of task-relevant ...

arXiv cs.RO ·Meiqi Wang, Shichao Li ·
compartilhar: