Paper Overview
Field: Computer Vision (CV) Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang Published: 2026-08-22 arXiv: 2608.20336
Abstract (English Translation)
Identity-preserving image generation becomes increasingly unreliable when scenes require multiple specified individuals. Beyond retaining each identity, the model must also bind each reference to distinct people and locations. This paper proposes WithEveryone, a unified framework for generating group images with up to ten reference identities.
WithEveryone:
- Injects each selected identity as an addressable token
- Predicts a structured identity-layout plan
- Renders the plan as visual conditioning
- Face similarity: improved from 0.462 (GPT-Image-2) to 0.499
- Copy-paste artifacts: reduced from 0.169 to 0.055
- Identity coverage: 97.3% of requested identities
- Duplication rate: only 2.8%
- Paper: https://arxiv.org/abs/2608.20336
Its core objective, Layout-Grounded ID Loss, directly supervises target identities using annotated facial regions, avoiding unstable embedding-based face matching. ID Representation Forcing trains a prediction for each identity before image synthesis.
Results
On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity: