> Paper: Reducing Pretraining-Generation Mismatch in Diffusion Language Models > Authors: Xiaocheng Lu, Huabin Liu, Song Guo, Jianguo Li > arXiv: https://arxiv.org/abs/2608.09424
An Overlooked Crack
If you have trained an autoregressive language model (the GPT kind), you probably know an unwritten rule: the input distribution at training time must match the input distribution at inference time.
Specifically, at inference the model receives a clean prompt and generates token by token. Training works the same way—the model sees a clean left context and predicts the next token. Training and inference share the same "clean prefix" interface.
This alignment seems so obvious that we rarely notice it exists.
But diffusion language models (dLLMs) break this assumption.
Diffusion LMs: The Price of Parallel Decoding
First, a recap of what diffusion language models do.
An autoregressive model generating 1,000 tokens needs 1,000 forward passes. A diffusion language model instead randomly masks all tokens first, then recovers them in parallel over multiple denoising rounds. In theory, it can generate hundreds of tokens in a few dozen steps—a clear speed advantage.
LLaDA, MDLM, and Plaid are representative works in this direction. In 2025, Ant Group's LLaDA scaled diffusion language models to 100B parameters, proving the path is engineering-feasible.
But dLLMs have an overlooked crack: the masking scheme at pretraining differs from the masking scheme at inference.
Where Exactly the Crack Is
In standard dLLM pretraining, masking is globally random—given a text sequence, some tokens are randomly selected and masked, and the model learns to recover them. This means:
- Prompt tokens may be masked
- Continuation tokens may also be masked
- Prompt and continuation are treated symmetrically
- Prompt: autoregressive loss (next-token prediction)
- Continuation: diffusion loss (recovering from masks)
- Average across 6 benchmarks: baseline → PCD, a 4.2% relative improvement (+2.56 points absolute)
- This is a continued-pretraining experiment on the existing LLaDA2-Mini
- Main mechanism comparison: 14.2% relative improvement (+4.86 points absolute)
- This is a from-scratch training experiment
- Standard dLLM pretraining performs well on the "random mask recovery" task (training loss drops, recovery quality improves).
- But that training task is not the inference task. Metrics on the training task mask true inference performance.
- Epanorthosis: standard loss masks the systematic emergence of AI-flavored rhetorical devices.
- Token Budget: standard loss masks the bimodal fate of CoT reasoning.
- QuantiBias: standard safety checks mask 24–27% bias introduced by quantization.
- PCD: standard dLLM training loss masks the training-inference interface mismatch.
But what about inference? At inference, the model receives a completely clean prompt; only the continuation needs to be masked and recovered.
The paper calls this inconsistency the Pretraining-Generation Mismatch:
| | Pretraining | Inference | |---|---|---| | Prompt state | Randomly masked | Fully clean | | Continuation state | Randomly masked | Fully masked → progressively denoised | | Prompt–continuation relation | Symmetric (equally masked) | Asymmetric (prompt is the condition) |
The consequence: the model never learns the task of "recovering the continuation conditioned on a clean prompt" during pretraining. It learns a more general but more diffuse task—"recovering randomly masked tokens in a randomly masked context."
An analogy: it's like a basketball player who practices shooting from any position in any posture, but games only require free throws from the free-throw line. When training and game distributions diverge, performance at the line suffers no matter how much you practice.
PCD: Aligning the Training Interface to the Inference Interface
The paper's solution is called PCD (Prefix-Conditioned Diffusion). The core idea is straightforward:
Treat the prompt as clean during training, and apply diffusion only to the continuation.
Concretely, PCD modifies three things:
1. Attention Mask
In standard dLLM pretraining, all tokens attend to each other (bidirectional attention). PCD makes the prompt part causal—continuation tokens can see all prompt tokens, but prompt tokens only see preceding prompt tokens.
This makes the prompt representations closer to an autoregressive model—clean, ordered, unpolluted by the continuation.
2. Corruption Mask
Standard dLLMs randomly mask the whole sequence. PCD masks only the continuation; the prompt stays clean throughout.
3. Label Construction
Standard dLLMs compute the loss over all tokens. PCD splits it into two parts:
Together, these three changes align the training "input interface" with the inference "input interface." During training, the model practices exactly what it must do at inference: recover the continuation given a clean prompt.
A Key Design: No New Components
PCD has an engineering-critical property: it needs no autoregressive decoder, no verifier, and no new inference mode.
What does this mean? PCD is a pure training-objective modification. At inference, a PCD-trained model runs the exact same pipeline as a standard dLLM—no inference code changes, no extra models, no new decoding strategy.
This matters because many training improvements demand corresponding inference changes (speculative decoding needs a draft model; Medusa needs multi-head decoding). PCD keeps all changes in training: zero inference cost.
The paper also runs a finer ablation: intra-sample prefix conditioning vs. inter-sample objective mixing. The former means "within one sample, prompt clean, continuation diffused"; the latter means "mixing AR samples and diffusion samples within a batch." Experiments show the former is the main contributor; the latter is an optional tuning knob.
This ablation is valuable because it tells us where the alignment signal comes from—not from batch-level mixing, but from sample-level interface alignment.
Experimental Results
The paper evaluates on two backbone models:
LLaDA2-Mini:
Qwen-1.7B:
Gains are consistent across both backbones, suggesting PCD is not a model-specific trick but a general improvement to the dLLM training objective.
Why This Direction Matters
Diffusion language models are among the most promising challengers to the autoregressive paradigm, for three reasons:
1. Inference speed: parallel decoding can generate hundreds of tokens in a few dozen steps—roughly an order of magnitude faster than autoregression. 2. Global planning: diffusion models see the global structure of the whole sequence while generating, theoretically better suited for long-text planning. 3. Controllability: the diffusion framework naturally supports infilling, editing, and conditional generation.
But dLLMs have long had a "good but not quite good enough" problem—close to autoregressive models on standard benchmarks, yet always slightly behind. This paper's diagnosis: part of that gap comes from the training-generation mismatch.
PCD patches part of the crack. The 4.2% and 14.2% gains are not earth-shattering, but they point to a systematic improvement direction: aligning the training interface with the inference interface.
Echoes of the "Evaluation Blind-Spot Law"
This paper can also be understood through the lens of an "evaluation blind-spot law":
The structure is isomorphic to several earlier papers:
An Honest Assessment
A few caveats are worth noting:
The gains are modest. 4.2% and 14.2% are relative improvements; the absolute gains are 2.56 and 4.86 points. Averaged over 6 benchmarks, that is roughly 0.4–0.8 absolute points per benchmark—a meaningful but not disruptive improvement.
Only two backbones validated. LLaDA2-Mini and Qwen-1.7B are relatively small models. Whether PCD holds at larger scale (e.g., 7B+) needs further verification.
No open-source code. The paper provides no GitHub link, which is an obstacle to reproduction and community follow-up.
PCD is continued pretraining, not from scratch, for LLaDA2-Mini. This keeps PCD's cost relatively low, while the Qwen-1.7B experiment trains from scratch at higher cost. Gains appear in both settings, indicating PCD is effective at different training stages.
Conclusion
The paper's core insight in one sentence:
Diffusion language models' training and inference objectives are misaligned—training randomly masks all tokens, while inference masks only the continuation. By aligning the training interface with the inference interface, PCD recovers 4–14% of performance without changing the inference pipeline.
The value of this insight lies not in the specific numbers, but in revealing a neglected design dimension: training-inference interface alignment. Autoregressive models are aligned for free (both train and inference predict the next token given a clean prefix), so the issue never surfaces in that paradigm. But once you switch to diffusion, alignment is no longer a free lunch—it must be explicitly designed.
A wake-up call for anyone building non-autoregressive models: is your training task really the same as your inference task?
---
Paper link: https://arxiv.org/abs/2608.09424