论文概要
研究领域: ML
作者: Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
发布时间: 2026-09-28
arXiv: 2609.35763
中文摘要
分布训练通过在冻结的表示空间中匹配真实与生成特征,为单步视觉生成提供集体监督。我们引入一个统一的理论框架,将分布建模与匹配散度分离,并通过 Wasserstein 梯度流将全局目标与逐点特征更新联系起来。在该框架下,FD-Loss 和高斯核漂移分别通过高斯最优传输和基于核密度的 KL 匹配被重新推导。该框架启发了 MGFlow——用高斯混合在全局矩和基于样本的表示之间以可调粒度建模特征分布。MGFlow 支持最优传输和基于分数的匹配,并将质量约束的样本分配与成对分量更新耦合,以解决仅靠混合表达力无法解决的模态崩溃问题。在 ImageNet 256×256 上,MGFlow 大幅超越 FD-Loss 基线,在 pMF-H 上取得 1.45 FDr⁶ 的 SOTA 结果,在 JiT-H 上取得 1.64。对于文本到图像生成,MGFlow 将 FLUX.2 [klein] 4B 后训练为单步生成器,在 GenEval 和 PickScore 上均超过原始四步模型。
原文摘要
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples m...
自动采集于 2026-09-30
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。