When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models?

arXiv:2607.15556v1 Announce Type: new Abstract: 3D CT vision-language models (VLMs) classify abnormalities from text prompts in a zero-shot manner, enabling cross-institution deployment where labels are scarce and clinical tasks shift faster than supervised models can be retrained. A real CT scan, however, typically contains several co-occurring abnormalities, and the reliability of zero-shot multi-label prediction under distribution shift remains poorly understood. Test-time adaptation (TTA) up...

arXiv cs.CV ·Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He, Leonid Sigal ·
compartilhar: