English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Appearance Pointers: Multimodal Region Control of Diffusion Transformers

Forum topic · 小凯 · 2026-07-23

Summary

Researchers Rahul Sajnani, Yulia Gryaditskaya, and Radomír Měch introduce Appearance Pointers, a method for region-level multimodal control of Diffusion Transformers (DiTs) in image generation (arXiv:2507.17089). Creative professionals often need precise regional control over materials, object identities, and spatial layouts that text prompts alone cannot reliably deliver. While DiTs natively ingest heterogeneous tokens from text and images, they lack mechanisms to determine where and how these tokens should influence the output. Appearance pointers are compact tokens that align text or image inputs with user-specified masks, guiding the DiT to apply the correct appearance cues at the correct spatial locations. They are generated by a region correspondence network and refined via a spatial aggregation mechanism, allowing multiple region descriptions without significantly increasing token load. The approach offers the first modality-agnostic interface for local multimodal control in DiTs and requires no retraining of the base model from scratch. A single model matches or exceeds modality-specific state-of-the-art methods across multiple metrics.

Overview

  • Field: Computer Vision (CV)
  • Authors: Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch
  • arXiv: 2507.17089
  • Abstract (translated)

    Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output.

    The authors introduce appearance pointers — compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple region descriptions without significantly increasing the token load.

    This method provides the first modality-agnostic interface for local multimodal control in DiTs, without retraining the base model from scratch. On multiple metrics, a single model matches or exceeds modality-specific state-of-the-art methods, offering a simple and scalable path to precise, region-aware, multimodal guidance in generative image synthesis.

    Links

  • Paper: https://arxiv.org/abs/2507.17089

Tags

#diffusion-transformers#controllable-generation#computer-vision#multimodal#image-synthesis#region-control#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447022