Summary
This paper introduces a method for fine-grained identity tuning in text-to-image personalization models. Generating and editing a person's face requires high precision, since even minor modifications can significantly alter perceived identity, yet existing personalization and editing methods built on general-purpose text-to-image models often lack this precision. Rather than editing a given image directly, the proposed identity-tuning approach modifies the latent representation of a specific identity, allowing the generation of diverse images that consistently depict the same edited identity. The authors explore the latent space of a pre-trained, frozen encoder for personalization, which consists of a set of latent tokens capturing different aspects of identity, often corresponding to specific spatial or semantic facial regions. They show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically consistent edits. Both qualitative and quantitative experiments demonstrate the ability to perform diverse local facial edits while preserving identity consistency across images. Authors: Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik. arXiv: 2607.11885.
Overview
- Field: Computer Vision (CV)
- Authors: Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik
- arXiv: 2607.11885
Summary
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits.
This paper presents a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity.
Approach
To enable fine-grained latent identity tuning, the authors explore the latent space of a pre-trained, frozen encoder for text-to-image personalization, leveraging its existing architecture to discover latent semantic directions. Key observations:
- The latent space consists of a set of latent tokens, each capturing a different aspect of identity.
- These tokens often correspond to specific spatial or semantic facial regions.
- Meaningful directions can be identified within this space and within subspaces defined by selected tokens.
This enables
localized, fine-grained, and semantically consistent facial edits.
Results
Both qualitative and quantitative experiments validate that the method can perform diverse local facial edits while preserving identity consistency across generated images.
---
Source: auto-collected forum post, original arXiv page: https://arxiv.org/abs/2607.11885
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178395142