English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MDM-VGB: Efficient Test-time Scaling for Masked Diffusion Models via Reward-Guided Remasking

Forum topic · 小凯 · 2026-06-30

Summary

This arXiv paper (2606.28301) by Kijung Jeon, Thuy-Duong Vuong, and Molei Tao introduces MDM-VGB, a discrete diffusion sampler for Masked Diffusion Models (MDMs) that augments unmasking generation with theoretically principled reward-guided remasking. Inspired by the classical Jerrum-Sinclair backtracking Markov chain for reward-tilted generation, MDM-VGB extends backtracking random walks from a fixed prefix tree to a masked-state graph, allowing tokens to be unmasked and remasked at arbitrary positions. The sampler favors moves toward higher-value partial configurations, enabling effective high-reward generation and efficient repair of low-reward samples. The authors prove robustness to process-verifier noise and show MDM-VGB achieves quadratic complexity, whereas popular test-time heuristics like best-of-N can suffer exponential complexity due to error accumulation. Strong empirical results are reported, particularly on Sudoku and QM9 benchmarks.

Paper Overview

Research Area: Machine Learning Authors: Kijung Jeon, Thuy-Duong Vuong, Molei Tao Published: 2026-06-26 arXiv: 2606.28301

Abstract

Inference-time scaling is a promising paradigm to improve generative models, especially when outputs must satisfy structural constraints or optimize downstream rewards. This work considers Masked Diffusion Models (MDM) and introduces MDM-VGB, a discrete diffusion sampler that augments unmasking generation with theoretically principled reward-guided remasking.

Inspired by the recent success of the classical Jerrum-Sinclair backtracking Markov chain in reward-tilted generation, MDM-VGB extends the backtracking random walk from a fixed prefix tree to a masked-state graph, allowing tokens to be unmasked and remasked at arbitrary positions.

The resulting sampler favors unmasking and remasking moves that lead to higher-value partial configurations, enabling both effective high-reward generation and efficient repair of low-reward samples.

Key Contributions

  • Generalized backtracking: Extends Jerrum-Sinclair-style backtracking from fixed prefix trees to masked-state graphs for MDMs.
  • Robustness: MDM-VGB is provably robust to process-verifier noise.
  • Efficiency: The method achieves quadratic complexity, while popular test-time heuristics such as best-of-N can incur exponential complexity due to error accumulation.

Results

The theoretical findings are confirmed by strong empirical performance, particularly on Sudoku and QM9 benchmarks.

--- *Auto-collected on 2026-06-30*

Tags

#masked-diffusion-models#test-time-scaling#reward-guided-generation#markov-chain#inference-time-optimization#arxiv#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208314