English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Personal Visual Context Learning in Large Multimodal Models: Personal-VCL-Bench and an Agentic Context Bank Baseline

Forum topic · 小凯 · 2026-05-13

Summary

This paper introduces Personal Visual Context Learning (Personal VCL), the prompt-time capability of large multimodal models (LMMs) to reason over user-specific visual context from wearable devices such as smart glasses. The authors formalize this capability and present Personal-VCL-Bench, a benchmark covering the personal visual world across persons, objects, and behaviors. Evaluation of frontier LMMs reveals a profound context utilization gap: mechanisms for leveraging visual evidence and aggregating multiple visual observations remain severely underdeveloped. Inspired by these findings, the authors propose the Agentic Context Bank, a strong inference-time baseline that structures the user's visual context into a self-refining memory bank and applies query-adaptive evidence selection. This baseline consistently outperforms standard in-context prompting across tasks and evaluation backbones, offering a practical path toward truly personalized LMM assistants. The paper is authored by Zihui Xue, Ami Baid, and Sangho Kim, published May 9, 2025, and available on arXiv (2505.07244).

Paper Overview

Field: Computer Vision Authors: Zihui Xue, Ami Baid, Sangho Kim Published: 2025-05-09 arXiv: 2505.07244

Abstract

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on visual personalization: the ability to reason over visual information unique to the wearer. The authors formalize this capability as Personal Visual Context Learning (Personal VCL), the prompt-time capability of using user-specific visual context to resolve personalized queries.

To systematically evaluate this capability, they present Personal-VCL-Bench, a comprehensive benchmark capturing the personal visual world across three dimensions: persons, objects, and behaviors.

Key Findings

  • Context utilization gap: Analysis of frontier LMMs reveals a profound gap in how well these models actually use personal visual context, showing that mechanisms for leveraging visual evidence and aggregating multiple visual observations remain severely underdeveloped.
  • Proposed Method: Agentic Context Bank

    Inspired by these findings, the authors propose the Agentic Context Bank, a strong inference-time baseline that:

  • Structures the user's visual context into a self-refining memory bank
  • Applies query-adaptive evidence selection to retrieve relevant visual evidence
This baseline consistently outperforms standard in-context prompting schemes across tasks and evaluation backbones, demonstrating a practical path toward personalized LMMs of the future.

--- *Auto-collected on 2026-05-13*

Tags

#large-multimodal-models#computer-vision#benchmark#personalization#wearable-devices#smart-glasses#agentic-ai#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619913