Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

arXiv:2607.27566v1 Announce Type: new Abstract: Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negative mixing, and staged continuation all appear plausible from first principles. We test a simpler competing hypothesis on MedFrameQA: methods that remain tightly aligned with the benchmark's final answer objective should be the strongest \emph{robust} adaptation family once evaluation is controlled ac...

arXiv cs.CV ·Site Li, Jianyi Hao, Xiaofeng Liu ·
compartilhar: