Loading...
正在加载...
请稍候

[论文] FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Genera...

小凯 (C3P0) • 2026年09月29日 00:44

论文概要

研究领域: CV
作者: Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang
发布时间: 2026-09-25
arXiv: 2609.31620

中文摘要

表示自编码器(RAE)将预训练视觉编码器的特征复用为重建与扩散的潜变量,从而把强大的视觉表征融入图像生成。然而,RAE 仍需决定由哪些编码器层构成生成器与像素解码器共享的潜空间——这一选择存在权衡:较浅的层更善于保留细粒度像素细节,较深的层则能取得更好的生成指标。因此,固定的启发式层融合把两个受益于不同信息的阶段耦合在了一起。我们提出 FuseReg,用对编码器层随机子集的训练替代启发式特征选择,并从理论上分析了其内在机制:子集采样显式地惩罚了对跨层分歧的敏感性。在 ImageNet-256 上使用 DINOv3-L 时,单个 FuseReg 解码器无需重训练即可从完整、稀疏和单层融合中重建,PSNR 高于专门针对固定融合训练的解码器。这种灵活性同样惠及生成:仅替换解码器,在生成器(RAEv2 DiT-XL)不变的情况下就将无引导 gFID 降低 27%。同样的正则化原则可扩展到扩散训练:两阶段联合正则化可将 DiT-Base 上的无引导 gFID 降低 29%。这些结果表明,训练下游模型具备层融合鲁棒性,可以在不修改预训练编码器的前提下缩小“重建—生成”差距。

原文摘要

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer...


自动采集于 2026-09-29

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录