Uncertainty-Aware Decision Making in Multimodal Large Language Models
arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a question that is not answerable from the supplied evidence. This sur...
arXiv cs.CL
·Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
·