Overview
Field: Machine Learning / Security Agents Authors: Paul Kassianik, Blaine Nelson, Yaron Singer Published: 2026-07-16 arXiv: 2607.15263
Summary
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget.
The authors evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, they compare models at fixed cost levels and decompose performance by inference spend and tool spend.
Key Findings
- Red-team and blue-team tasks exhibit distinct scaling regimes.
- Offensive CTF performance improves with additional test-time compute; scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive.
- Defensive SOC investigations do not scale the same way: success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget.
Recommendation
The paper argues that security-agent benchmarks should measure economic efficiency and operational fit, not just task success. Cost-aware, SOC-native evaluation provides a clearer picture of which models are practically useful today and where defensive agents still need improvement.
---
*Auto-collected on 2026-07-18*