Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

arXiv:2608.17306v1 Announce Type: new Abstract: While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen classes, leading to a rapid degradation in performance against adversarial examples of unseen classes as training progresses. We empirically identify that this degradation stems from the tendency of the model to learn pseudo-robust features (i.e., non-generalizable shortcuts)...

arXiv cs.CV ·Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen, Hua Liu, Yu Zhang ·
compartilhar: