DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
arXiv:2607.02551v1 Announce Type: new Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current models often fail to find the change and provide reliable evidence. We propose DELTAVID, a verifiable proxy-task framework that enhances fine-grained spatiotemporal percept...
arXiv cs.CV
·Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, Shuo Yang
·