English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation with Up to 10 Reference Identities

Forum topic · 小凯 · 2026-08-24

Summary

WithEveryone is a unified framework for identity-preserving group image generation that supports up to ten reference identities in a single scene. The method injects each selected identity as an addressable token, predicts a structured identity-layout plan, and renders that plan as visual conditioning. Its key objective, a layout-grounded ID loss, directly supervises the intended identity using annotated face regions, avoiding unstable embedding-based face matching among multiple noisy predicted faces. On identity-disjoint benchmarks, WithEveryone achieves the highest target-context identity similarity, improving face similarity from GPT-Image-2's 0.462 to 0.499 while reducing copy-paste artifacts from 0.169 to 0.055. It covers 97.3% of requested identities with a repetition rate of only 2.8%. The results show that explicit identity-layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference face copying. Paper: arXiv 2608.20336.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
  • arXiv: 2608.20336
  • Abstract

    Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces.

    The authors introduce WithEveryone, a unified framework for generating group images containing up to ten reference identities. WithEveryone:

  • Injects each selected identity as an addressable token
  • Predicts a structured identity-layout plan
  • Renders that plan as visual conditioning
  • Its key objective, a layout-grounded ID loss, uses annotated face regions to directly supervise the intended identity, avoiding unstable embedding-based face matching. The ID representation enforces a per-identity prediction before image synthesis.

    Results

    On identity-disjoint benchmarks, WithEveryone achieves:

  • Highest target-context identity similarity, improving face similarity from 0.462 (GPT-Image-2) to 0.499
  • Copy-paste artifacts reduced from 0.169 to 0.055
  • Coverage of 97.3% of requested identities with a repetition rate of only 2.8%
These results indicate that explicit identity-layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference face copying.

---

*Auto-collected on 2026-08-24.*

Tags

#witheveryone#identity-preserving-generation#group-image-generation#text-to-image#computer-vision#diffusion-models#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633915