Paper Overview
Field: Computer Vision Authors: Wenti Yin, Xiang Wang, Huaxin Zhang arXiv: 2508.05149
Abstract
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality.
To this end, the authors propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through interaction between competing explanations and video evidence.
Method
1. Contrast construction: Using public safety knowledge, CEAVAD builds hazard-normal event contrasts, pairing each hazardous mechanism with a generic normal explanation and a mechanism-specific benign counterpart. 2. Contrastive adjudication: It judges whether the target interval better supports the hazardous explanation or its benign competing explanation, generating a correctable contrastive boundary proposal for the target. 3. Evidence-based verdict: CEAVAD adjudicates between the competing explanations, determining whether the hazard hypothesis can withstand the video evidence, supporting both temporally localized anomaly detection and evidence-based explanation.
Results
Experiments on three widely used VAD benchmarks show that CEAVAD achieves state-of-the-art performance under the training-free paradigm.
---
*Auto-collected on 2026-08-12.*