[论文] dQwen3.5: Hybrid-Attention Diffusion Language Models

研究领域: NLP 作者: Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai 发布时间: 2026-09-17 arXiv: 2609.20751

论文概要

研究领域: NLP 作者: Anton Xue, Litu Rout, Aditya Akella, Adam Klivans, Sujay Sanghavi, Sanjay Shakkottai 发布时间: 2026-09-17 arXiv: 2609.20751

中文摘要

把预训练自回归(AR)模型改造为扩散语言模型(DLM)是一条低成本路线。几乎所有此类改造都以全注意力 Transformer 为起点,但 AR 建模已转向混合架构——交替使用注意力层与 RNN 层。这给改造带来障碍:与注意力不同,RNN 在结构上是因果的,难以双向化。尽管存在这一错配,本文仍研究此类骨干能否成为有效的 DLM:我们对 0.8B、2B、4B、9B 四个规模的 Qwen3.5 进行改造,得到 dQwen3.5 家族。我们发现混合骨干可以成为高效的改造起点:与全注意力对照组相比,混合骨干用约一半的 token 量即可达到给定训练损失。在各规模上,dQwen3.5 展现出与全注意力 DLM 相似的任意顺序解码行为,并在并行解码下表现强劲。

原文摘要

Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attentio...


*自动采集于 2026-09-20*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens