论文概要
研究领域: CV
作者: Kaixiong Gong, Xin Cai, Bin Lin, Hao Wang, Yunlong Lin, Mingzhe Zheng, Bohao Li, Jian-Wei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Xiangyu Yue
发布时间: 2026-07-24
arXiv: 2607.22531
中文摘要
统一多模态模型寻求一个共享的视觉token空间,以同时支持多模态理解和图像生成。离散方法通过共享码本统一接口,而连续流程通常依赖两种不同表示——用于理解的语义特征(如ViT)和用于合成的低级隐变量(如VAE)——导致隐空间不匹配。本文提出Twins,一种统一的连续token空间,通过在相同token网格上按通道拼接ViT和VAE特征形成,因此序列长度不变且注意力成本不增加。然而,在Diffusion Transformer中联合建模Twins暴露出严重的优化不平衡:模型能很好地拟合ViT组件,但难以匹配VAE隐变量分布。本文将这种不平衡追溯到三个异质性来源:频率偏差、内在维度和条件对齐与条件独立的不确定性。为解决这一问题,本文将焦点回归目标适配于流匹配,对误差较大的VAE维度进行加权,更好地平衡ViT和VAE组件之间的优化。在ImageNet上,这相比朴素MSE损失在无需分类器自由引导的情况下获得了高达10.57的gFID提升。Twins在多模态理解基准测试上也具有竞争力,并提高了重建保真度,缩小了面向理解和面向生成的表示之间的差距。
原文摘要
Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations -- semantic features (e.g., ViT) for understanding and low-level latents (e.g., VAE) for synthesis -- resulting in mismatched latent spaces. We propose Twins, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase. However, jointly modeling Twins in a Diffusion Transformer exposes a severe optimization imbalance: the model fits the ViT component well but struggles to match the VAE latent distribution. We t...
自动采集于 2026-07-28
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。