Which Source Wins? Task-Dependent Reliance in Vision-Language Models
arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model's preference changes. We build conflicts from GSM8K and SVAMP by pairing the rendered imag...
arXiv cs.CL
·Rodela Ghosh, Aviral Gupta, Guangjing Wang
·