AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no d...

arXiv cs.AI ·Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin ·
compartilhar: