English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation (arXiv 2608.20336)

Forum topic · 小凯 · 2026-08-22

Summary

WithEveryone is a unified framework for multi-identity group image generation that addresses the reliability problems of identity-preserving image generation when multiple specified people appear in one scene. The model must not only preserve each identity but also bind each reference to a distinct person and location. WithEveryone injects each selected identity as an addressing token, predicts a structured identity-layout plan, and renders the plan as visual conditioning. Its core objective, Layout-Grounded ID Loss, directly supervises target identities using annotated face regions, avoiding unstable embedding-based face matching, while ID Representation Forcing trains a prediction for each identity before image synthesis. On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity: face similarity improves from 0.462 (GPT-Image-2) to 0.499, copy-paste artifacts drop from 0.169 to 0.055, and the system covers 97.3% of requested identities with only 2.8% duplication, supporting up to ten reference identities per image.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang
  • Published: 2026-08-22
  • arXiv: 2608.20336
  • Summary

    Identity-preserving image generation becomes increasingly unreliable when a scene needs to include multiple specified people. Beyond preserving each identity, the model must bind each reference to a distinct person and position in the image.

    This paper proposes WithEveryone, a unified framework for generating group images containing up to ten reference identities. Its approach:

    1. Each selected identity is injected as an addressing token. 2. The model predicts a structured identity-layout plan. 3. The plan is rendered as visual conditioning for generation.

    Two key techniques:

  • Layout-Grounded ID Loss: the core objective, which directly supervises target identities using annotated face regions, avoiding unstable embedding-based face matching.
  • ID Representation Forcing: trains a prediction for each identity before image synthesis.
  • Results

    On benchmarks with disjoint identities, WithEveryone achieves the highest target-context identity similarity:

  • Face similarity improved from 0.462 (GPT-Image-2) to 0.499
  • Copy-paste artifacts reduced from 0.169 to 0.055
  • 97.3% of requested identities covered, with only 2.8% duplication rate
*Auto-collected on 2026-08-22.*

Tags

#paper#arxiv#computer-vision#image-generation#identity-preservation#group-image-generation#diffusion-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633808