English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive LLM Security Agents

Forum topic · 小凯 · 2026-07-18

Summary

This paper by Paul Kassianik, Blaine Nelson, and Yaron Singer (arXiv:2607.15263) argues that security-agent evaluations should measure economic efficiency, not just peak offensive capability under generous inference budgets. The authors evaluate language-model security agents through a cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, they compare models at fixed cost levels and decompose performance by inference spend and tool spend. Results reveal distinct scaling regimes: offensive CTF performance improves with additional test-time compute, and scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigations do not scale the same way—success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget. The authors advocate cost-aware, SOC-native benchmarks that capture operational fitness and economic efficiency alongside task success, clarifying which models are practically useful today and where defensive agents still need improvement.

Paper Overview

Field: Machine Learning Authors: Paul Kassianik, Blaine Nelson, Yaron Singer Published: 2026-07-16 arXiv: 2607.15263

Abstract (Translated)

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry query, and enrichment request consumes budget. We evaluate language-model security agents through this cost-success lens on offensive Cybench challenges and defensive Splunk BOTS v1 investigation challenges. Instead of reporting only best-case success, we compare models at fixed cost levels and decompose performance by inference spend and tool spend.

Our results show distinct scaling regimes for red- and blue-team tasks. Offensive CTF performance improves with additional test-time compute; scaled open-weight models can approach frontier proprietary systems while remaining cost-competitive. Defensive SOC investigations do not scale the same way: success depends more on disciplined tool use, telemetry navigation, and selective enrichment than on raw reasoning budget alone. We argue that security-agent benchmarks should measure economic efficiency and operational fitness alongside task success. Cost-aware, SOC-native evaluation provides a clearer picture of which models are actually useful today and where defensive agents still need to improve.

Key Takeaways

  • Success rate alone is an incomplete metric for security agents; budget consumption per step matters in operations.
  • Offensive (CTF) capability scales with test-time compute; open-weight models can rival frontier systems at competitive cost.
  • Defensive (SOC investigation) success hinges on tool discipline and telemetry navigation, not just more inference spend.
  • The paper proposes cost-aware, SOC-native benchmarking as a better lens for evaluating security agents.

Tags

#machine-learning#security-agents#llm-evaluation#cybersecurity#cost-aware-benchmarks#ctf#soc#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178433593