English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Forum topic · 小凯 · 2026-07-18

Summary

This paper (arXiv:2607.15263) by Paul Kassianik, Blaine Nelson, and Yaron Singer argues that security-agent evaluations should go beyond peak success rates measured under generous inference budgets. Since operational security work consumes budget at every reasoning step, tool call, telemetry query, and enrichment request, the authors evaluate language-model security agents through a cost-success lens on offensive Cybench CTF challenges and defensive Splunk BOTS v1 SOC investigation challenges. Rather than reporting best-case success, they compare models at fixed cost levels and decompose performance into inference spend and tool spend. Results reveal distinct scaling regimes: offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigations do not scale the same way—success depends more on disciplined tool use, telemetry navigation, and selective enrichment than raw reasoning budget. The authors advocate benchmarks that measure economic efficiency and operational fit alongside task success.

Overview

Field: Machine Learning / Security Agents Authors: Paul Kassianik, Blaine Nelson, Yaron Singer Published: 2026-07-16 arXiv: 2607.15263

Summary

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget.

The authors evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, they compare models at fixed cost levels and decompose performance by inference spend and tool spend.

Key Findings

  • Red-team and blue-team tasks exhibit distinct scaling regimes.
  • Offensive CTF performance improves with additional test-time compute; scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive.
  • Defensive SOC investigations do not scale the same way: success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget.

Recommendation

The paper argues that security-agent benchmarks should measure economic efficiency and operational fit, not just task success. Cost-aware, SOC-native evaluation provides a clearer picture of which models are practically useful today and where defensive agents still need improvement.

---

*Auto-collected on 2026-07-18*

Tags

#security-agents#llm-evaluation#cost-aware-benchmarks#cybench#soc-investigation#ctf#red-team#blue-team

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433584