PE-Field 4D: Video Generation Models as Canvas

arXiv:2607.15667v1 Announce Type: new Abstract: Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of positional encoding in video diffusion transformers and show that it provides a useful spatial bias for geometry-aware control. Specifically, if reference tokens are encoded according to their projected locations in the target view, th...

arXiv cs.CV ·Yunpeng Bai, Haoxiang Li, Qixing Huang ·
compartilhar: