English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation in Autonomous Driving

Forum topic · 小凯 · 2026-07-28

Summary

CARA (Concept-Aware Risk Attention) is an intrinsically interpretable spatio-temporal framework for collision anticipation in autonomous driving, presented in arXiv paper 2607.22494. Addressing the opacity of feature-driven models and the low fidelity of post-hoc explanations, CARA derives domain-grounded risk concepts from accident narratives, aligns them with video frames via vision-language similarity, and organizes them into evolving concept trajectories. These trajectories serve as explicit, dynamic risk evidence that directly guides spatial attention, temporal attention, and anticipation, coupling interpretability tightly with the prediction process rather than treating it as an afterthought. Experiments on three benchmarks show CARA consistently improves anticipation accuracy and warning lead time over strong baselines while providing sparse, semantically grounded concept evidence, making it suitable for safety-critical autonomous driving applications.

Overview

  • Field: Computer Vision (CV)
  • Authors: Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, Yue Su, Jinbo Su, Yi Hong
  • Published: 2026-07-24
  • arXiv: 2607.22494
  • Abstract

    Collision anticipation in autonomous driving requires not only accurate early warnings but also interpretable reasoning about what risk factors are being tracked and how risk evolves over time. Existing methods fall short in this regard: feature-driven models are opaque, post-hoc explanations often lack fidelity, and concept-based methods are mostly designed for static recognition rather than dynamic driving scenes.

    Key Contributions

    The authors propose CARA (Concept-Aware Risk Attention), an intrinsically interpretable spatio-temporal framework for collision anticipation:

  • Derives domain-grounded risk concepts from accident narratives
  • Aligns these concepts with video frames via vision-language similarity
  • Organizes them into evolving concept trajectories
These trajectories provide explicit risk evidence that guides spatial attention, temporal attention, and anticipation, allowing semantic concepts to directly influence where the model attends and how it predicts risk over time. By treating semantic risk factors as dynamic intermediate evidence rather than auxiliary post-hoc explanations, CARA couples interpretability tightly with the prediction process.

Results

Extensive experiments on three benchmarks demonstrate that CARA consistently improves anticipation accuracy and warning lead time, outperforming strong baselines while providing sparse and semantically grounded concept evidence.

Tags

#autonomous-driving#collision-anticipation#interpretability#vision-language#computer-vision#arxiv#risk-attention

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503745