Paper Overview
Field: Computer Vision Authors: Wenti Yin, Xiang Wang, Huaxin Zhang Published: 2026-08-11 arXiv: 2508.03800
Introduction
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality.
Proposed Method: CEAVAD
The authors propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing explanations and video evidence.
The method works in three stages:
1. Hazard-benign event contrast construction: Using public safety knowledge, each hazardous mechanism is paired with a generic normal description and a mechanism-specific benign counterpart. 2. Contrastive boundary proposal: For a target interval, CEAVAD determines whether the hazardous explanation or its benign competitor is better supported by the evidence, yielding a correctable contrastive boundary proposal. 3. Adjudication: CEAVAD adjudicates among competing explanations to decide whether the hazard hypothesis withstands scrutiny against the video evidence, supporting both temporally localized anomaly detection and evidence-based explanations.
Results
Experiments on three widely used VAD benchmarks demonstrate that CEAVAD achieves state-of-the-art performance under the training-free paradigm.
Links
- arXiv: https://arxiv.org/abs/2508.03800
*Auto-collected on 2026-08-12*