RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
arXiv:2607.00310v1 Announce Type: new Abstract: Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains. We study parameter-efficient adaptation of a pretrained foundation video world model to retail scenes: when synchronized egocentric and exocentric video of the same activity are available, which viewpoint of training data produces the s...
arXiv cs.CV
·Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi
·
// relacionados
Leia também
Editorial
RynnValue: o relógio do vídeo como recompensa para robôs
Blog
Anthropic aplica marca d'água a todas as saídas do Claude globalmente, com marcas que "podem persistir mesmo após alguma edição"
Blog
webAI lança TwIL-LM: uma família de modelos de lógica formal de 1,7B e 3B para autoformalização em hardware local
Blog