Paper Overview
Research Area: Machine Learning Authors: Kijung Jeon, Thuy-Duong Vuong, Molei Tao Published: 2026-06-26 arXiv: 2606.28301
Abstract
Inference-time scaling is a promising paradigm to improve generative models, especially when outputs must satisfy structural constraints or optimize downstream rewards. This work considers Masked Diffusion Models (MDM) and introduces MDM-VGB, a discrete diffusion sampler that augments unmasking generation with theoretically principled reward-guided remasking.
Inspired by the recent success of the classical Jerrum-Sinclair backtracking Markov chain in reward-tilted generation, MDM-VGB extends the backtracking random walk from a fixed prefix tree to a masked-state graph, allowing tokens to be unmasked and remasked at arbitrary positions.
The resulting sampler favors unmasking and remasking moves that lead to higher-value partial configurations, enabling both effective high-reward generation and efficient repair of low-reward samples.
Key Contributions
- Generalized backtracking: Extends Jerrum-Sinclair-style backtracking from fixed prefix trees to masked-state graphs for MDMs.
- Robustness: MDM-VGB is provably robust to process-verifier noise.
- Efficiency: The method achieves quadratic complexity, while popular test-time heuristics such as best-of-N can incur exponential complexity due to error accumulation.
Results
The theoretical findings are confirmed by strong empirical performance, particularly on Sudoku and QM9 benchmarks.
--- *Auto-collected on 2026-06-30*