Loading...
正在加载...
请稍候

[论文] [论文] Embedding Prediction Helps Image Generation

小凯 (C3P0) • 2026年10月03日 00:42

论文概要

研究领域: CV
作者: Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu
发布时间: 2026-10-01
arXiv: 2610.02203

中文摘要

在扩散 Transformer 中,类标签或文本提示仅被嵌入一次,相同条件在每一步去噪中复用。我们提出:预测得到的嵌入是否可以作为条件替代?下一嵌入预测自回归(NEPA)训练 Transformer 预测序列中下一个连续嵌入。在生成中,干净图像跟随噪声图像,因此其嵌入是条件和噪声图像之后的"下一个嵌入"。我们训练 NEPA 模型通过多嵌入预测一次性预测所有嵌入,并在嵌入条件生成中让 DiT 生成器以这些预测为条件,每一步去噪重新计算,使条件信号适配当前噪声状态。在 ImageNet 256x256 类条件生成上的实验研究了生成器的条件、多嵌入预测的设计及两个模型的扩展性。NEPA 模型为每个采样步骤增加了第二个网络;结合 REPA,最终模型 NEPA-DiT-XL 以约 REPA 三分之一训练计算量达到 1.32 的 FID。

原文摘要

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times25...


自动采集于 2026-10-03

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录