ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
arXiv:2607.28751v1 Announce Type: new Abstract: Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding or add computation through explicit rationale tokens and latent autoregressive states. Although token expansion can improve complex matching, serial generation increases retrieval latency and makes the final embedding depend on generated intermediate states. This raises a...
arXiv cs.CV
·Shijie Wang, Xiangzhao Hao, Yueti Li, Guangyu Cao, Xinyu Tang, Haiyun Guo
·
// relacionados
Leia também
Blog
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Editorial
CAPA: o benchmark que mede se o assistente de código aprende com você — ou repete a mesma pergunta
Blog
Why biological data matters more in AI drug discovery
Blog