Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
arXiv:2607.02588v1 Announce Type: new Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build higher-level memory stores organized around temporal proximity rather than explicit causal links, leaving multi-hop narrative reasoning to be reconstructed by the LLM at ev...
arXiv cs.CV
·Yixin Ji, Fanghua Ye, Juntao Li, Bo Zhao, Zexuan Qiu, Zhaopeng Tu, Liefeng Bo, Min Zhang
·
// relacionados
Leia também
Blog
Um Guia de Programação para a Programação de GPU Baseada em Tiles da NVIDIA: De cuTile e Kernels Triton até Flash Attention
Blog
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
Blog
Grupos terroristas estão usando todos os principais chatbots de IA para planejamento de ataques e desenvolvimento de armas
Blog