DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
arXiv:2606.27499v1 Announce Type: new Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down. We introduce DMV-Bench (Code: https://github.com/yyyujintang/DMV-Bench), the first interactive benchmark for multimodal-agent visual memory. DMV-Bench is built on a controlled home-furnishing e-commerce catalogue of ...
arXiv cs.CV
·Yujin Tang, Chenming Shang, Ruize Xu, Nikhil Singh
·
// relacionados
Leia também
Blog
Um Guia de Programação para a Programação de GPU Baseada em Tiles da NVIDIA: De cuTile e Kernels Triton até Flash Attention
Blog
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
Blog
Grupos terroristas estão usando todos os principais chatbots de IA para planejamento de ataques e desenvolvimento de armas
Blog