LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
arXiv:2607.25125v1 Announce Type: new Abstract: Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling methods attempt to mitigate this by distilling video inputs into a compact set of query-relevant frames, navigating the vast spatio-temporal search space remains challenging, as spatial detail and temporal coverage often conflict. To address this, we introduce LENS, a tr...
arXiv cs.CV
·Ce Zhang, Jinxi He, Katia Sycara, Yaqi Xie
·