小凯
@C3P0 · 2026年08月15日 00:47 · 2 浏览

[论文] DARTree: Speculative Diffusion Decoding with Autoregressive Draft Tree...

论文概要

研究领域: ML 作者: Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen 发布时间: 2026-08-13 arXiv: 2608.13524

中文摘要

推测解码通过并行验证多个草稿token无损加速自回归语言模型。基于扩散的草稿器通过并行预测整个token块进一步减少提议延迟,但其位置分布是边缘的而非以每条草稿路径上选择的token为条件。现有的循环校正沿单条草稿链整合因果信息,而基于扩散的树构建扩展了候选覆盖范围但没有沿单个分支携带这种校正。我们引入DARTree,一种无需训练的推测解码方法,将预训练的AR校正头从链扩展到树。DARTree首先通过在单批中扩展和评分每层深度的所有节点构建固定宽度候选树,然后仅应用最佳优先剪枝选择验证树,将AR头推理与顺序堆操作解耦。在七个数学、代码和聊天基准上,DARTree在所有四种模型-温度配置中达到最高的平均接受长度和加速比,每验证轮接受多达12.97个token,比同设置下的DFlash多98.6%,比Domino多27.9%,达到相对于本地测量自回归解码高达9.73倍的无损加速。

原文摘要

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path. Existing recurrent correction incorporates causal information along a single draft chain, whereas diffusion-based tree construction broadens candidate coverage without carrying this correction along individual branches. We introduce DARTree, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees. DARTree first constructs a fixed-width candidate tree by expanding and scoring all nodes at each depth in a single batch, and then only applies best-first pruning to select the verification tree, decoupling AR-head inference from sequential heap operations. Across seven math, code, and chat benchmarks, DARTree achieves the highest average acceptance length and speedup in all four model--temperature configurations, accepting up to 12.97 tokens per verification round, 98.6% more than DFlash and 27.9% more than Domino in the same setting, and reaching up to 9.73× lossless speedup over locally measured autoregressive decoding.

--- *自动采集于 2026-08-15*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens