Paper Overview
Field: NLP Authors: Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli Posted: 2026-06-25 arXiv: 2606.19224
Abstract (translated)
Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer — a baseline reliability property called for by emerging AI evaluation guidelines. The authors introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs. A Bayesian item-response model separates ordering noise from per-facet bias, and a same-ordering control estimates the decoder-stochastic floor for observed flips.
Key Findings
- No model is order-invariant: none of the 18 audited MLLMs passes the reliability check. Screened per-facet panel-mean flip rates span 24–50%.
- Beyond decoder noise: a Gemini same-ordering control at temperature 0 estimates a substantial ordering excess over the same-input decoder-noise floor, ruling out sampling randomness as the cause.
- Capability helps but doesn't solve it: stronger capability predicts fewer flips, but the best models still flip on 13.4% of trials.
- Prompt mitigation is modality-conditional: training-free prompt changes tested on Gemini do not transfer from text to visual reasoning.
Implications
These results suggest that prompt-level mitigation alone is unlikely to deliver general order robustness, motivating future work on training-time and architectural approaches. The authors propose cross-ordering flip rate as a standard reporting axis for MLLM evaluation.
--- *Auto-collected on 2026-06-26*