ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
ABACUS is a unified vision-language model that performs object counting and related tasks through innovative spatial grounding, boundary-aware counting policies, and self-critical…
Hugging Face · Daily Papers
·Anindya Mondal, Sauradip Nag
·
·▲ 4 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Anindya Mondal, Sauradip Nag, Anjan Dutta
- 4 upvotes da comunidade
- Temas: vision-language model, object counting, crowd counting, referring-expression counting, count-faithful image generation, unified foundation model
Resumo
Resumo original (em inglês), extraído do paper:
ABACUS is a unified vision-language model that performs object counting and related tasks through innovative spatial grounding, boundary-aware counting policies, and self-critical learning strategies.