English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection (CEAVAD)

Forum topic · 小凯 · 2026-08-11

Summary

This paper introduces CEAVAD (Contrastive Event Adjudication for training-free Video Anomaly Detection), a new approach to identifying and temporally localizing abnormal events in videos without any training on target-domain labels. The authors argue that prior training-free methods rely on pretrained models' semantic knowledge to describe hazards, but richer descriptions capture hazard resemblance without resolving abnormality. CEAVAD reframes the inference unit from isolated anomaly concepts to falsifiable event hypotheses, establishing an explanatory boundary through competition between interpretations and video evidence. It builds hazard-vs-benign event contrasts using public-safety knowledge, pairs each hazard mechanism with a generic normal description and a mechanism-specific benign counterpart, and adjudicates between competing explanations to determine whether a hazard hypothesis withstands the video evidence. Experiments on three widely used VAD benchmarks show CEAVAD achieves state-of-the-art performance in the training-free paradigm while supporting both temporally localized anomaly detection and evidence-based explanations.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Wenti Yin, Xiang Wang, Huaxin Zhang
  • Published: 2026-08-11
  • arXiv: 2508.03800
  • Summary

    Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality.

    To address this, the authors propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing interpretations and video evidence.

    Key Points

  • Problem reframing: Move from asking "does this look like an anomaly?" to "does the hazard hypothesis withstand the video evidence when compared to a benign alternative?"
  • Hazard-benign contrasts: Public-safety knowledge is used to construct contrasts pairing each hazard mechanism with (1) a generic normal description and (2) a mechanism-specific benign counterpart.
  • Contrastive boundary proposal: For a target interval, the method determines whether the hazard interpretation or its benign competitor is better supported by the visual evidence, yielding a revisable contrastive boundary proposal.
  • Event adjudication: Final decision adjudicates between competing interpretations, supporting both temporally localized anomaly detection and evidence-based explanations.
  • Results: Experiments on three widely used VAD benchmarks show that CEAVAD achieves state-of-the-art performance in the training-free paradigm.

Original Abstract (excerpt)

> Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality. To this end, we propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary thr...

*Source: arXiv 2508.03800, auto-collected 2026-08-12.*

Tags

#video-anomaly-detection#training-free#contrastive-learning#event-adjudication#computer-vision#arxiv-2508-03800#temporal-localization#foundation-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633362