[论文] Discrete Beckmann Transport Models for One-Step Language Modeling and ...
研究领域: ML 作者: Sophia Tang, Shiyi Wang 发布时间: 2026-09-14 arXiv: 2609.15903
论文概要
研究领域: ML 作者: Sophia Tang, Shiyi Wang 发布时间: 2026-09-14 arXiv: 2609.15903
中文摘要
离散扩散和流模型是自回归语言模型的有前景替代方案,但将多步采样压缩为更少步骤通常需要蒸馏预训练教师模型。这使学生受限于教师的质量,并需要昂贵的两阶段训练流水线。我们引入离散Beckmann传输模型(DBTM),构建于时间无关流上,其自治传输映射可证明地将环境空间中的任何点在单步内传送到单纯形顶点的固定点。我们证明该固定点性质由守恒方程表征,其残差可直接从数据最小化,从而消除了对教师流和时间条件的需求。在此构造下,部分训练的映射对应于有限时间截断的流,因此生成归结为迭代一个映射直至达到固定点。我们进一步将映射扩展为部分上下文插值器,使额外的函数评估充当细化步骤而非ODE积分步骤。在语言建模和推理任务上,DBTM实现了一步和少步生成,质量和准确率均优于离散扩散和连续流基线。
原文摘要
Discrete diffusion and flow models are a promising alternative to autoregressive language models, but compressing many-step sampling into fewer steps typically requires distilling a pretrained teacher model. This caps the student at the teacher's quality and requires a costly two-stage training pipeline. We introduce Discrete Beckmann Transport Models (DBTM), built on a time-independent flow whose autonomous transport map provably carries any point in the ambient space to a fixed point on the vertices of the simplex in a single step. We show that this fixed-point property is characterized by a conservation equation whose residual can be minimized directly from data, removing the requirement for a teacher flow and time conditioning. Under this construction, a partially trained map correspon...
*自动采集于 2026-09-16*
#论文 #arXiv #ML #小凯