Summary
WithEveryone is a unified framework for multi-identity group image generation that addresses the reliability problems of identity-preserving image generation when multiple specified people appear in one scene. The model must not only preserve each identity but also bind each reference to a distinct person and location. WithEveryone injects each selected identity as an addressing token, predicts a structured identity-layout plan, and renders the plan as visual conditioning. Its core objective, Layout-Grounded ID Loss, directly supervises target identities using annotated face regions, avoiding unstable embedding-based face matching, while ID Representation Forcing trains a prediction for each identity before image synthesis. On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity: face similarity improves from 0.462 (GPT-Image-2) to 0.499, copy-paste artifacts drop from 0.169 to 0.055, and the system covers 97.3% of requested identities with only 2.8% duplication, supporting up to ten reference identities per image.
Paper Overview
- Research area: Computer Vision (CV)
- Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang
- Published: 2026-08-22
- arXiv: 2608.20336
Summary
Identity-preserving image generation becomes increasingly unreliable when a scene needs to include multiple specified people. Beyond preserving each identity, the model must bind each reference to a distinct person and position in the image.
This paper proposes WithEveryone, a unified framework for generating group images containing up to ten reference identities. Its approach:
1. Each selected identity is injected as an addressing token.
2. The model predicts a structured identity-layout plan.
3. The plan is rendered as visual conditioning for generation.
Two key techniques:
- Layout-Grounded ID Loss: the core objective, which directly supervises target identities using annotated face regions, avoiding unstable embedding-based face matching.
- ID Representation Forcing: trains a prediction for each identity before image synthesis.
Results
On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity:
- Face similarity improved from 0.462 (GPT-Image-2) to 0.499
- Copy-paste artifacts reduced from 0.169 to 0.055
- 97.3% of requested identities covered, with only 2.8% duplication rate
*Auto-collected on 2026-08-22.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633808