[论文] [论文] Sphere Encoder 2
研究领域: CV 作者: Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein 发布时间: 2026-10-01 arXiv: 2610.02208
论文概要
研究领域: CV 作者: Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein 发布时间: 2026-10-01 arXiv: 2610.02208
中文摘要
Sphere Encoder 是一种自编码器,通过从高维隐空间球面上解码随机点来生成图像。我们发现原始表述存在两个限制生成质量的缺陷。其一,相对于极点,随机点在编码后的隐空间分布集中于赤道附近,但训练时的旋转从未到达该区域,留下了一个限制一步生成的缺口。其二,使用逐像素重建损失进行生成训练会鼓励解码器对合理图像取平均,产生缺乏高频细节的模糊图像。我们提出 Sphere Encoder 2 同时解决这两个问题,在保持自编码器速度和简洁性的同时大幅提升图像生成质量。模型已开源。
原文摘要
Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.c...
*自动采集于 2026-10-03*
#论文 #arXiv #CV #小凯