ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a unified vision-language model that performs object counting and related tasks through innovative spatial grounding, boundary-aware counting policies, and self-critical…

Hugging Face · Daily Papers ·Anindya Mondal, Sauradip Nag · ·▲ 4 upvotes

Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.

Autores: Anindya Mondal, Sauradip Nag, Anjan Dutta

  • 4 upvotes da comunidade
  • Temas: vision-language model, object counting, crowd counting, referring-expression counting, count-faithful image generation, unified foundation model

Resumo

Resumo original (em inglês), extraído do paper:

ABACUS is a unified vision-language model that performs object counting and related tasks through innovative spatial grounding, boundary-aware counting policies, and self-critical learning strategies.

Onde ler

compartilhar: