CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

arXiv:2607.25294v1 Announce Type: new Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted repor...

arXiv cs.CV ·Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang ·
compartilhar: