LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

arXiv:2607.25125v1 Announce Type: new Abstract: Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling methods attempt to mitigate this by distilling video inputs into a compact set of query-relevant frames, navigating the vast spatio-temporal search space remains challenging, as spatial detail and temporal coverage often conflict. To address this, we introduce LENS, a tr...

arXiv cs.CV ·Ce Zhang, Jinxi He, Katia Sycara, Yaqi Xie ·
compartilhar: