[论文] Hierarchical Continuous Diffusion Language Models

研究领域: NLP 作者: Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing 发布时间: 2026-10-01 arXiv: 2610.02193

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing 发布时间: 2026-10-01 arXiv: 2610.02193

中文摘要

离散扩散语言模型为需要双向推理与全局约束满足的任务,提供了有吸引力的自回归替代方案。但它们共享一个结构性瓶颈:并行解码时,每个 token 从自身边缘分布独立采样,切断了同批解码 token 之间的统计依赖。连续扩散语言模型通过去噪共享连续状态规避了这一点,但去噪器只看到该状态,直到最终解码前没有任何东西将其约束到合法的 token 配置。为此我们提出层次化连续扩散语言模型(HC-DLM):将离散 token 生成与连续隐轨迹耦合进单一、有原理的去噪过程,训练目标派生自 token 似然的变分下界。与近期把连续上下文附加到自洽离散链上的方法不同,HC-DLM 让隐变量成为唯一持久的生成状态——token 每步从中读出,并作为下一步隐变量更新的脚手架反馈回去。在结构化推理(数独)、数学规划(Countdown)与语言建模(LM1B)上,同等模型规模下 HC-DLM 均优于离散与连续扩散基线:数独与 Countdown 的谜题准确率更高,LM1B 生成困惑度更低。项目页面:https://hc-dlm.github.io/

原文摘要

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training o...


*自动采集于 2026-10-04*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens