Paper Overview
Field: cs.AI Authors: Divyanshu Kumar, Rohith HN, Nitin Aravind Birur, Sahil Agarwal, Prashanth Harshangi Published: 2026-09-13 arXiv: 2609.11030
Abstract (Original)
AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog containing records of agent-related events disclosed over multiple years. Each record includes supporting evidence, a stable identifier, and missingness-aware labels for causal role, disclosure class, mechanism, and outcome. Among the generative-system records in which the agent acted, a substantial portion involved realized harm. Realized outcomes concentrate in in-the-wild and safety-failure records, while responsible disclosures and research demonstrations are overwhelmingly demonstrated; the aggregate share therefore characterizes collection composition rather than deployment risk. After initial curation, a second human reviewer checked all records and their existing labels for completeness and correctness. In a deployment-analogue audit, InjecAgent's cases occupy three of AIR's twelve surfaces and are all attacker-triggered, whereas AIR contains no-adversary safety failures as well. AIR supports source-grounded case retrieval and evaluation-scope auditing, not failure-rate or control-efficacy estimation.
Key Takeaways
- Purpose: General incident repositories lack the mechanism-level detail needed to compare real-world AI agent failures with the coverage of agent-security evaluations. AIR fills this gap with a source-linked, structured catalog.
- Record structure: Each entry includes supporting evidence, a stable identifier, and missingness-aware labels for causal role, disclosure class, mechanism, and outcome.
- Quality control: After initial curation, a second human reviewer independently verified all records and labels for completeness and correctness.
- Key finding: Realized harm clusters in in-the-wild and safety-failure records; responsible disclosures and research demonstrations are largely demonstrated-only, so the overall harm share reflects collection composition, not deployment risk.
- Evaluation audit: InjecAgent's benchmark cases map to only three of AIR's twelve failure surfaces and are all attacker-triggered, while AIR additionally includes safety failures involving no adversary — highlighting gaps in current security benchmarks.
- Intended use: Source-grounded case retrieval and evaluation-scope auditing; explicitly not for estimating failure rates or control efficacy.
*Auto-collected on 2026-09-13*