[论文] Appearance Pointers -- Multimodal Region Control of Diffusion Transfor...
论文概要
研究领域: CV 作者: Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch 发布时间: 2026-07-22 arXiv: 2507.17089
中文摘要
可控图像生成对创意专业人士而言仍具挑战性,他们通常需要对材质、物体身份和空间布局进行精确的区域控制,而这无法仅通过文本提示可靠实现。扩散Transformer(DiT)可以原生处理来自文本和图像的异构token,但缺乏确定这些token应在何处、如何影响输出的机制。我们引入\textbf{外观指针}(appearance pointers),这是一种紧凑token,通过将文本或图像输入与用户指定掩码对齐,引导DiT在正确的空间位置关注正确的外观线索。外观指针由区域对应网络生成,并通过空间聚合机制细化,使模型能够处理多个区域描述而不会显著增加token负载。我们的方法首次为DiT中的局部多模态控制提供了模态无关的接口,且无需从头重新训练基础模型。在多项指标上,我们的单模型达到或超越了模态特定的最先进方法,为生成图像合成中精确、区域感知、多模态引导提供了一条简单且可扩展的路径。
原文摘要
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regi...
--- *自动采集于 2026-07-23*
#论文 #arXiv #CV #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens