论文概要
研究领域: CV
作者: Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Divyansh Srivastava, Bingnan Li, Zhuowen Tu
发布时间: 2026-10-07
arXiv: 2610.10497
中文摘要
我们提出QuadTok,一个新颖的视觉token化和自回归图像生成框架。与传统方法使用2D网格或1D token序列不同,我们提出层次化四叉树结构,在2D空间绑定和1D序列灵活性之间架起桥梁。QuadTok分词器动态地将表示容量分配给视觉复杂区域,同时让同质区域保持粗分辨率。与固定256 token网格相比,我们的ImageNet训练分词器在ImageNet上节省约10%的token,在零样本迁移到COCO数据集时节省9%,同时保持相当的重建保真度。此外,树结构引入的自然因果性无缝支持自回归图像生成。在给定四叉树拓扑的条件下,我们9.47亿参数的GPT式生成模型在ImageNet 256×256基准上实现了2.08 gFID。此外,利用四叉树结构保持的强空间相关性,QuadTok生成器实现了零样本空间可控图像生成能力。代码:https://github.com/myc634/QuadTok
原文摘要
We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. C...
自动采集于 2026-10-09
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。