Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign cri...

arXiv cs.CV ·Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li ·
compartilhar: