English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Facet-Probe: Auditing Order Sensitivity in Multimodal Large Language Models

Forum topic · 小凯 · 2026-06-26

Summary

A new paper introduces Facet-Probe, a five-facet audit measuring how sensitive multimodal large language models (MLLMs) are to order-irrelevant permutations of inputs. The audit covers option ordering, evidence-chunk ordering, document ranking, image-set ordering, and mixed-modality ordering across 18 frontier and open-weight MLLMs. A Bayesian item-response model separates ordering noise from per-facet bias, while a same-ordering control estimates the decoder-stochastic floor. Key finding: none of the 18 audited MLLMs is order-invariant; screened per-facet panel-mean flip rates range from 24% to 50%. A Gemini control at temperature 0 shows a substantial ordering excess over the decoder-noise floor, indicating genuine ordering sensitivity rather than sampling randomness. Capability reduces but does not eliminate flips—the best models still flip answers on 13.4% of trials. Training-free prompt-based mitigations tested on Gemini proved modality-conditional and did not transfer from text to visual reasoning. The authors conclude that prompt-level mitigation alone is unlikely to deliver general order robustness, motivating training-time and architectural approaches, and propose cross-ordering flip rate as a standard reporting axis for MLLM evaluation. Paper: arXiv 2606.19224 by Akshay Paruchuri, Sanmi Koyejo, and Ehsan Adeli.

Paper Overview

Field: NLP Authors: Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli Posted: 2026-06-25 arXiv: 2606.19224

Abstract (translated)

Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer — a baseline reliability property called for by emerging AI evaluation guidelines. The authors introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs. A Bayesian item-response model separates ordering noise from per-facet bias, and a same-ordering control estimates the decoder-stochastic floor for observed flips.

Key Findings

  • No model is order-invariant: none of the 18 audited MLLMs passes the reliability check. Screened per-facet panel-mean flip rates span 24–50%.
  • Beyond decoder noise: a Gemini same-ordering control at temperature 0 estimates a substantial ordering excess over the same-input decoder-noise floor, ruling out sampling randomness as the cause.
  • Capability helps but doesn't solve it: stronger capability predicts fewer flips, but the best models still flip on 13.4% of trials.
  • Prompt mitigation is modality-conditional: training-free prompt changes tested on Gemini do not transfer from text to visual reasoning.

Implications

These results suggest that prompt-level mitigation alone is unlikely to deliver general order robustness, motivating future work on training-time and architectural approaches. The authors propose cross-ordering flip rate as a standard reporting axis for MLLM evaluation.

--- *Auto-collected on 2026-06-26*

Tags

#multimodal-llms#evaluation#order-sensitivity#benchmarking#robustness#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208137