Visual Distribution Anchoring for Efficient Prompt Tuning
arXiv:2607.28967v1 Announce Type: new Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch. We propose VDA (Visual Distribution Anchoring), a training-free target adaptation framework that augments a frozen semantic classifier with class-level visual...
arXiv cs.CV
·Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi
·