English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HOMIE: Human-Object Centric Video Personalization via Multimodal Integration

Forum topic · 小凯 · 2026-07-22

Summary

HOMIE is a new framework for human-object centric video personalization (HOCVP), a core task in subject-driven video generation. Existing methods face two key limitations: they struggle to balance high subject fidelity with accurate interaction patterns between humans and diverse objects—especially when objects represent abstract concepts like logos—and they lack mechanisms to exploit intra-subject references such as OCR maps or multi-view inputs. HOMIE addresses both challenges in a unified framework that handles inter-subject and intra-subject input settings together. Compared with prior approaches, HOMIE introduces an improved strategy for integrating multimodal large language models (MLLMs) to extract reference-level relational knowledge, without compromising the controllability of the text encoder or requiring costly re-alignment. The method combines global multimodal guidance with modality-reference embeddings to achieve state-of-the-art performance. The paper (arXiv:2607.18217, cs.CV) is authored by Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, and 6 other researchers.

Paper Overview

  • Field: Computer Vision (cs.CV)
  • Authors: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, et al. (11 authors total)
  • Published: 2026-07-20
  • arXiv: 2607.18217
  • Background

    Human-object centric video personalization (HOCVP) is a central task in subject-driven video generation. Existing approaches suffer from two key limitations:

    1. Most methods focused on inter-subject personalization struggle to balance high subject fidelity with accurate interaction patterns between a person and diverse objects—particularly when the objects represent abstract concepts such as logos. 2. Although intra-subject references (e.g., OCR maps, multi-view inputs) can enhance subject fidelity, most existing work lacks mechanisms to understand this latent correspondence.

    Proposed Method

    HOMIE is a unified HOCVP framework that handles both inter-subject and intra-subject input settings. Compared with previous methods, HOMIE proposes:

  • A better MLLM integration strategy to extract reference-level relational knowledge, without impairing the text encoder's controllability or incurring expensive re-alignment
  • Global multimodal guidance
  • Modality-reference embedding
  • Together, these components deliver state-of-the-art performance.

    Links

  • Paper: https://arxiv.org/abs/2607.18217

Tags

#computer-vision#video-generation#personalization#multimodal#mllm#subject-driven-generation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447001