Summary
HOMIE is a new framework for human-object centric video personalization (HOCVP), a core task in subject-driven video generation. Existing methods face two key limitations: they struggle to balance high subject fidelity with accurate interaction patterns between humans and diverse objects—especially when objects represent abstract concepts like logos—and they lack mechanisms to exploit intra-subject references such as OCR maps or multi-view inputs. HOMIE addresses both challenges in a unified framework that handles inter-subject and intra-subject input settings together. Compared with prior approaches, HOMIE introduces an improved strategy for integrating multimodal large language models (MLLMs) to extract reference-level relational knowledge, without compromising the controllability of the text encoder or requiring costly re-alignment. The method combines global multimodal guidance with modality-reference embeddings to achieve state-of-the-art performance. The paper (arXiv:2607.18217, cs.CV) is authored by Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, and 6 other researchers.
Paper Overview
- Field: Computer Vision (cs.CV)
- Authors: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, et al. (11 authors total)
- Published: 2026-07-20
- arXiv: 2607.18217
Background
Human-object centric video personalization (HOCVP) is a central task in subject-driven video generation. Existing approaches suffer from two key limitations:
1. Most methods focused on inter-subject personalization struggle to balance high subject fidelity with accurate interaction patterns between a person and diverse objects—particularly when the objects represent abstract concepts such as logos.
2. Although intra-subject references (e.g., OCR maps, multi-view inputs) can enhance subject fidelity, most existing work lacks mechanisms to understand this latent correspondence.
Proposed Method
HOMIE is a unified HOCVP framework that handles both inter-subject and intra-subject input settings. Compared with previous methods, HOMIE proposes:
- A better MLLM integration strategy to extract reference-level relational knowledge, without impairing the text encoder's controllability or incurring expensive re-alignment
- Global multimodal guidance
- Modality-reference embedding
Together, these components deliver state-of-the-art performance.
Links
- Paper: https://arxiv.org/abs/2607.18217
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178447001