[论文] [论文] StableVQ: Practical Guidelines for Stable Vector-Quantized T...

论文概要 研究领域: CV 作者: Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu et al. 发布时间: 2026-09-22 arXiv: 2609.26774

论文概要

研究领域: CV 作者: Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu et al. 发布时间: 2026-09-22 arXiv: 2609.26774

中文摘要

向量量化(VQ)是支撑现代自回归与掩码图像生成模型的离散视觉分词器基础。近期共享投影码本方法虽大幅提升码本利用率,训练稳定性仍是关键且探索不足的难题。我们认为根源在编码器-解码器与码本训练的纠缠:两者都无法可靠独立履职,系统只能在二者碰巧合作时运转——脆弱条件,恰在训练最受压力时崩溃。StableVQ 重新审视各模块的恰当学习目标:(1)Dynamic STE 纠正编码器目标的不稳定性,即使码本利用率低也能在离散正则下稳健优化重建空间;(2)Region VQ Loss 重构码本目标,使其独立保证对编码器输出分布的完全追踪,无需依赖编码器振荡驱动激活;(3)Decoupled Schedule 为两者分配独立学习率调度,匹配各自职责的不同优化动态。StableVQ 构建于共享投影码本之上,轻量且无新增可学习参数。ImageNet 实验显示在多样码本大小与初始化下,训练稳定性、码本利用率与重建质量均持续提升。

原文摘要

Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.


*自动采集于 2026-09-24*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens