Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models
Accepted to the Conference on Neural Information Processing Systems (NeurIPS 2026)
Diffusion-based multimodal large language models (dMLLMs) predict tokens at many masked positions in parallel. Their decoding quality depends not only on the confidence of each prediction, but also on whether the tokens committed together draw on complementary visual evidence.
Existing confidence-based decoding ranks masked positions independently and commits the top-K. In multimodal settings, several high-confidence tokens selected in the same step can attend to the same image regions. This visual redundancy wastes parallel decoding capacity and limits the quality of the generated response.
We introduce the Visual Redundancy Index (VRI) to measure this overlap and Visual-Redundancy-Controlled Decoding (VRCD) to control it. VRCD is a training-free, inference-time method that uses token-to-image attention to select visually complementary positions, directly improving decoding quality by reducing redundant visual grounding.
Across a range of multimodal benchmarks VRCD lowers both visual redundancy and remaining-position entropy at modest runtime cost. Under longer decoding budgets it yields relative accuracy gains of up to 18.8% on M3CoT and 6.9% on MMBench over confidence-based decoding.
