English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Forum topic · 小凯 · 2026-08-22

Summary

WithEveryone is a unified framework for identity-preserving group image generation featuring up to ten reference identities, presented in arXiv paper 2608.20336 by Hengyuan Xu, Qixun Wang, Yiji Cheng, and Miles Yang. The method injects each selected identity as an addressable token, predicts a structured identity-layout plan, and renders that plan as visual conditioning. A Layout-Grounded ID Loss directly supervises target identities using annotated facial regions, avoiding unstable embedding-based face matching, while ID Representation Forcing trains per-identity predictions before image synthesis. On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 (GPT-Image-2) to 0.499, reducing copy-paste artifacts from 0.169 to 0.055, covering 97.3% of requested identities with only a 2.8% duplication rate. Full paper: https://arxiv.org/abs/2608.20336.

Paper Overview

Field: Computer Vision (CV) Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang Published: 2026-08-22 arXiv: 2608.20336

Abstract (English Translation)

Identity-preserving image generation becomes increasingly unreliable when scenes require multiple specified individuals. Beyond retaining each identity, the model must also bind each reference to distinct people and locations. This paper proposes WithEveryone, a unified framework for generating group images with up to ten reference identities.

WithEveryone:

  • Injects each selected identity as an addressable token
  • Predicts a structured identity-layout plan
  • Renders the plan as visual conditioning
  • Its core objective, Layout-Grounded ID Loss, directly supervises target identities using annotated facial regions, avoiding unstable embedding-based face matching. ID Representation Forcing trains a prediction for each identity before image synthesis.

    Results

    On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity:

  • Face similarity: improved from 0.462 (GPT-Image-2) to 0.499
  • Copy-paste artifacts: reduced from 0.169 to 0.055
  • Identity coverage: 97.3% of requested identities
  • Duplication rate: only 2.8%
  • Links

  • Paper: https://arxiv.org/abs/2608.20336
--- *Auto-collected on 2026-08-22*

Tags

#witheveryone#image-generation#identity-preservation#computer-vision#arxiv#diffusion-models#multi-subject-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633787