PE-Field 4D: Video Generation Models as Canvas
arXiv:2607.15667v1 Announce Type: new Abstract: Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of positional encoding in video diffusion transformers and show that it provides a useful spatial bias for geometry-aware control. Specifically, if reference tokens are encoded according to their projected locations in the target view, th...
arXiv cs.CV
·Yunpeng Bai, Haoxiang Li, Qixing Huang
·
// relacionados
Leia também
Editorial
Wan-Dancer: o vídeo de dança que a difusão não conseguia manter por mais de 20 segundos
Blog
DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse
Blog
Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
Blog