Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification

arXiv:2607.15615v1 Announce Type: new Abstract: Vision-language models trained with contrastive objectives have shown promise in medical image analysis. However, conventional global image-text alignment is ill-suited for mammography, where diagnostically relevant lesions are spatially localized and occupy only a small fraction of the image. Subtle morphological cues critical for malignancy assessment can be diluted when representations are learned at the whole-image level. In this work, we propo...

arXiv cs.CV ·Zhengbo Zhou, Jiren Li, Dooman Arefan, Margarita Zuley, Shandong Wu ·
compartilhar: