Paper Overview
- Field: Computer Vision (CV)
- Authors: Wenti Yin, Xiang Wang, Huaxin Zhang
- Published: 2026-08-11
- arXiv: 2508.03800
- Problem reframing: Move from asking "does this look like an anomaly?" to "does the hazard hypothesis withstand the video evidence when compared to a benign alternative?"
- Hazard-benign contrasts: Public-safety knowledge is used to construct contrasts pairing each hazard mechanism with (1) a generic normal description and (2) a mechanism-specific benign counterpart.
- Contrastive boundary proposal: For a target interval, the method determines whether the hazard interpretation or its benign competitor is better supported by the visual evidence, yielding a revisable contrastive boundary proposal.
- Event adjudication: Final decision adjudicates between competing interpretations, supporting both temporally localized anomaly detection and evidence-based explanations.
- Results: Experiments on three widely used VAD benchmarks show that CEAVAD achieves state-of-the-art performance in the training-free paradigm.
Summary
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality.
To address this, the authors propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing interpretations and video evidence.
Key Points
Original Abstract (excerpt)
> Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality. To this end, we propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary thr...
*Source: arXiv 2508.03800, auto-collected 2026-08-12.*