Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign cri...
arXiv cs.CV
·Jiasheng Li, Zhong Ji, Yan Zhang, Huihui Li
·
// relacionados
Leia também
Blog
Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights
Blog
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Blog
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Blog