English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

TPIPS: A Text-Prompted Image Perceptual Similarity Metric (arXiv 2607.18237)

Forum topic · 小凯 · 2026-07-22

Summary

Human judgments of visual similarity are context-dependent—two images may be similar in shape but differ greatly in color—yet existing perceptual similarity metrics compress these nuances into a single scalar. Researchers including Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, and Eli Shechtman address this gap in arXiv paper 2607.18237 (cs.CV, cs.LG). They introduce a large-scale dataset of human similarity judgments based on image triplets, each annotated along multiple free-form semantic similarity dimensions. Extensive benchmarking shows that state-of-the-art vision-language models (VLMs) still lag significantly behind human annotator consensus. Fine-tuning VLMs on this data yields TPIPS (Text-Prompted Image Perceptual Similarity), a metric that captures multiple senses of visual similarity conditioned on a text prompt. TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. The authors also demonstrate new applications in text-guided retrieval, compositional search, and fine-grained evaluation of generative models.

Paper Overview

Field: Computer Vision (CV) Authors: Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, et al. (7 authors total) Published: 2026-07-20 arXiv: 2607.18237 Categories: cs.CV, cs.LG

Summary

Human judgments of visual similarity are context-dependent. For example, two images may be similar in shape but differ dramatically in color. However, existing perceptual similarity metrics compress these nuances into a single scalar value and cannot be conditioned on specific aspects of similarity.

To bridge this gap, the authors introduce a large-scale dataset of human similarity judgments, collected on image triplets, where each triplet is annotated along multiple free-form semantic similarity dimensions.

Extensive benchmarking of state-of-the-art vision-language models (VLMs) reveals a significant performance gap compared to human annotator consensus. Leveraging this data, the authors fine-tune VLMs to produce TPIPS (Text-Prompted Image Perceptual Similarity), a metric that captures multiple semantic senses of visual similarity conditioned on a specified text prompt.

Key Contributions

  • A large-scale human similarity judgment dataset built on image triplets with free-form, multi-dimensional annotations.
  • Benchmarking that exposes a substantial gap between frontier VLMs and human consensus on similarity judgments.
  • TPIPS, a fine-tuned, text-prompted perceptual similarity metric that aligns more closely with human perception and generalizes reliably beyond the training distribution.
  • Demonstrated new capabilities in text-guided retrieval, compositional search, and fine-grained evaluation of generative models.
--- *Auto-collected on 2026-07-22*

Tags

#computer-vision#perceptual-similarity#vision-language-models#tpips#machine-learning#generative-models#arxiv#dataset

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446994