Summary
WithEveryone is a unified framework for identity-preserving group image generation that supports up to ten reference identities in a single scene. The method injects each selected identity as an addressable token, predicts a structured identity-layout plan, and renders that plan as visual conditioning. Its key objective, a layout-grounded ID loss, directly supervises the intended identity using annotated face regions, avoiding unstable embedding-based face matching among multiple noisy predicted faces. On identity-disjoint benchmarks, WithEveryone achieves the highest target-context identity similarity, improving face similarity from GPT-Image-2's 0.462 to 0.499 while reducing copy-paste artifacts from 0.169 to 0.055. It covers 97.3% of requested identities with a repetition rate of only 2.8%. The results show that explicit identity-layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference face copying. Paper: arXiv 2608.20336.
Paper Overview
- Field: Computer Vision (CV)
- Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
- arXiv: 2608.20336
Abstract
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces.
The authors introduce WithEveryone, a unified framework for generating group images containing up to ten reference identities. WithEveryone:
- Injects each selected identity as an addressable token
- Predicts a structured identity-layout plan
- Renders that plan as visual conditioning
Its key objective, a layout-grounded ID loss, uses annotated face regions to directly supervise the intended identity, avoiding unstable embedding-based face matching. The ID representation enforces a per-identity prediction before image synthesis.
Results
On identity-disjoint benchmarks, WithEveryone achieves:
- Highest target-context identity similarity, improving face similarity from 0.462 (GPT-Image-2) to 0.499
- Copy-paste artifacts reduced from 0.169 to 0.055
- Coverage of 97.3% of requested identities with a repetition rate of only 2.8%
These results indicate that explicit identity-layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference face copying.
---
*Auto-collected on 2026-08-24.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633915