Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

arXiv:2607.29079v1 Announce Type: new Abstract: Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing Fast-dLLM outputs with the same model's unaccelerated outputs. Across the mild parallelism induced in our long-form setting (1.05--1.25 committed tokens per step), confidence-threshold tuning changes decoding behavior but...

arXiv cs.CL ·Yaoxuan Dou, Yang Shu ·
compartilhar: