MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook...
arXiv cs.CV
·Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang
·
// relacionados