English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WALDO: Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Forum topic · 小凯 · 2026-05-08

Summary

WALDO is a training-free framework that improves zero-shot anomaly localisation in medical imaging by reformulating it as a comparative inference problem against reference distributions of normal anatomy. Grounded in optimal transport theory, it uses entropy-weighted Sliced Wasserstein distances for anatomically-aware reference selection from DINOv2 patch distributions, Goldilocks zone sampling that exploits the non-monotonic relationship between reference similarity and localisation accuracy, and self-consistency aggregation via weighted non-maximum suppression. The authors theoretically analyse the Goldilocks effect through distributional divergence, showing that moderately similar references minimise a bias-variance trade-off. On the NOVA brain MRI benchmark, WALDO with Qwen2.5-VL-72B achieves 43.5 ± 1.6% mAP@30 (95% CI: [40.4, 46.7]), a 19% relative improvement over zero-shot baselines. Cross-model evaluation with GPT-4o and Qwen3-VL-32B shows consistent gains (32.0% mAP@30 for both), with paired McNemar tests confirming statistical significance (p<0.01).

Overview

Field: Computer Vision Authors: Bernhard Kainz, Johanna P Mueller, Matthew Baugh, Cosmin Bercea Published: 2026-05-06 arXiv: 2605.05161

Abstract

Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is fundamentally limited by the absence of healthy anatomical context. We reformulate zero-shot localisation as a comparative inference problem in which anomalies are identified through structured comparison against reference distributions of normal anatomy.

We introduce WALDO, a training-free framework grounded in optimal transport theory that enables comparative reasoning through:

1. Entropy-weighted Sliced Wasserstein distances for anatomically-aware reference selection from DINOv2 patch distributions 2. Goldilocks zone sampling exploiting the non-monotonic relationship between reference similarity and localisation accuracy 3. Self-consistency aggregation via weighted non-maximum suppression

We theoretically analyse the Goldilocks effect through distributional divergence, and show that references with moderate similarity minimize a bias-variance trade-off in comparative visual reasoning.

Results

On the NOVA brain MRI benchmark:

  • WALDO + Qwen2.5-VL-72B: 43.5 ± 1.6% mAP@30 (95% CI: [40.4, 46.7]), a 19% relative improvement over zero-shot baselines
  • GPT-4o: 32.0 ± 6.5% mAP@30
  • Qwen3-VL-32B: 32.0 ± 6.6% mAP@30
Paired McNemar tests confirm statistical significance (p<0.01).

Code

Source code is available at https://github.com/bkainz/WALDO_MICCAI26_demo.

--- *Auto-collected on 2026-05-08*

Tags

#out-of-distribution-detection#vision-language-models#medical-imaging#optimal-transport#anomaly-localisation#brain-mri#zero-shot#miccai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619596