FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
arXiv:2607.25266v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases sharply as video sequence length increases, and exposing the model to more irrelevant content measurably reduces its accuracy. In this paper, we address the problem of maximizing query-relevant information in a frame subset selected at inference time, without training. FO...
arXiv cs.CV
·Ghazal Kaviani, Ghassan AlRegib
·