论文概要
研究领域: NLP
作者: Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing
发布时间: 2026-10-01
arXiv: 2610.02193
中文摘要
离散扩散语言模型为需要双向推理与全局约束满足的任务,提供了有吸引力的自回归替代方案。但它们共享一个结构性瓶颈:并行解码时,每个 token 从自身边缘分布独立采样,切断了同批解码 token 之间的统计依赖。连续扩散语言模型通过去噪共享连续状态规避了这一点,但去噪器只看到该状态,直到最终解码前没有任何东西将其约束到合法的 token 配置。为此我们提出层次化连续扩散语言模型(HC-DLM):将离散 token 生成与连续隐轨迹耦合进单一、有原理的去噪过程,训练目标派生自 token 似然的变分下界。与近期把连续上下文附加到自洽离散链上的方法不同,HC-DLM 让隐变量成为唯一持久的生成状态——token 每步从中读出,并作为下一步隐变量更新的脚手架反馈回去。在结构化推理(数独)、数学规划(Countdown)与语言建模(LM1B)上,同等模型规模下 HC-DLM 均优于离散与连续扩散基线:数独与 Countdown 的谜题准确率更高,LM1B 生成困惑度更低。项目页面:https://hc-dlm.github.io/
原文摘要
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training o...
自动采集于 2026-10-04
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。