AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no d...
arXiv cs.AI
·Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
·
// relacionados
Leia também
Modelo
internlm/Intern-S2-Preview-397B
Blog
Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification
Blog
Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach
Blog