Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retr…
Hugging Face · Daily Papers
·Suhyeong Park, Junha Jung
·
·▲ 11 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang
- 11 upvotes da comunidade
- Temas: vision-language retrieval, late interaction, token compression, object-aware merging, post-projector tokens, centroid compression
Resumo
Resumo original (em inglês), extraído do paper:
Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.Onde ler
// relacionados
Leia também
Blog
Flight attendants freaked out that Google is buying tons of Spirit employee data
Blog
Anthropic says any lab can now let a language model agent run the whole protein design stack
Blog
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Blog