English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark

Forum topic · 小凯 · 2026-05-22

Summary

WikiVQABench is a human-curated, knowledge-grounded visual question answering (VQA) benchmark introduced by Basel Shbita, Pengyuan Li, and Anna Lisa Gentile (arXiv:2505.15981). Unlike traditional VQA benchmarks that emphasize perception-based tasks solvable from visual content alone, WikiVQABench requires external knowledge beyond what is visible in the image. The benchmark is built by systematically combining Wikipedia images, their associated article captions, and structured knowledge from Wikidata. A pipeline using large language models generates candidate multiple-choice image-question-answer sets, all of which are reviewed and curated by human annotators to ensure factual correctness, visual-text consistency, and that each question genuinely needs external knowledge in addition to visual evidence. The dataset includes a large collection of Wikipedia images and curated multiple-choice questions designed to benchmark knowledge-aware vision-language models (VLMs). Evaluations across 15 VLMs ranging from 256M to 90B parameters show a wide performance spread, with accuracy between 24.7% and 75.6%, demonstrating the benchmark's effectiveness in distinguishing model capabilities on knowledge-intensive reasoning. Dataset and benchmark code are publicly available.

Paper Overview

Research Area: Computer Vision (CV) Authors: Basel Shbita, Pengyuan Li, Anna Lisa Gentile Published: 2025-05-20 arXiv: 2505.15981

Abstract

Visual Question Answering (VQA) benchmarks have largely emphasized perception-based tasks that can be solved from visual content alone. In contrast, many real-world scenarios require external knowledge that is not directly observable in the image to answer correctly. The authors introduce WikiVQABench, a human-curated knowledge-grounded VQA benchmark constructed by systematically combining Wikipedia images, their associated article captions, and structured knowledge from Wikidata.

Key Points

  • Construction pipeline: Large language models (LLMs) generate candidate multiple-choice image-question-answer sets from Wikipedia images, captions, and Wikidata structured knowledge.
  • Human curation: All generated instances are reviewed and curated by human annotators to ensure factual correctness, visual-text consistency, and that each question requires external knowledge in addition to visual evidence.
  • Scale: The benchmark contains a large collection of Wikipedia images and curated multiple-choice questions, designed to evaluate knowledge-aware vision-language models (VLMs).
  • Evaluation results: Testing across 15 VLMs (256M–90B parameters) shows a wide performance range of 24.7%–75.6% accuracy, demonstrating the benchmark's effectiveness in differentiating models on knowledge-intensive reasoning.
  • Availability: The dataset and benchmark code are publicly available.

Significance

WikiVQABench addresses a gap in existing VQA evaluation: most benchmarks reward purely visual perception, while real-world applications demand reasoning that fuses vision with external world knowledge. This benchmark provides a rigorous testbed for measuring that fusion capability.

---

*Auto-collected on 2026-05-22.*

Tags

#wikivqabench#visual-question-answering#benchmark#vision-language-models#wikidata#wikipedia#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620576