Summary
OMIBench is a new benchmark for evaluating Olympiad-level reasoning in large vision-language models (LVLMs) when evidence is distributed across multiple images. Existing Olympiad-level multimodal benchmarks emphasize single-image analysis, neglecting cross-image contextual information. OMIBench addresses this gap with problems drawn from biology, chemistry, mathematics, and physics Olympiads, each accompanied by manually annotated rationales and evaluation protocols supporting both exact and semantic answer matching. Extensive experiments reveal substantial performance gaps among current models: even the strongest LVLMs, such as Gemini-3-Pro, achieve only about 50% on the benchmark. These results establish OMIBench as a focused resource for researching and improving multi-image reasoning capabilities in LVLMs. The paper, authored by Qiguang Chen, Chengyu Luan, and Jiajun Wu, is available on arXiv.
Paper Overview
Field: NLP
Authors: Qiguang Chen, Chengyu Luan, Jiajun Wu
Published: 2026-04-22
arXiv: 2604.20806
Abstract
Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level multimodal reasoning benchmarks for these models often emphasize single-image analysis and fail to exploit contextual information across multiple images. We present OMIBench, a benchmark designed to evaluate Olympiad-level reasoning when the required evidence is distributed over multiple images. It contains problems from biology, chemistry, mathematics, and physics Olympiads, together with manually annotated rationales and evaluation protocols for both exact and semantic answer matching.
Key Findings
- Across extensive experiments on OMIBench, meaningful performance gaps are observed in existing models.
- Even the strongest LVLMs, such as Gemini-3-Pro, attain only about 50% on the benchmark.
- These results position OMIBench as a focused resource for researching and improving multi-image reasoning in LVLMs.
Links
- arXiv: <https://arxiv.org/abs/2604.20806>
*Auto-collected on 2026-04-24*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177618693