Two VentureBeat Pulse Research surveys, both released July 16, 2026, together depict a harsh reality: AI agents are running faster than enterprises are prepared to handle.
Survey 1: The Agent Security Gap (107 enterprises)
Sample: 107 enterprises with 100+ employees, single June 2026 survey (self-selected sample — directional signal, not precise measurement).
| Key metric | Value | |---|---| | Experienced an agent security incident or near-miss | 54% (18% confirmed incidents + 36% near-misses) | | Per-agent scoped identity | Only 32% | | Agents sharing credentials (API keys / borrowed human credentials) | 69% (multi-select) | | High-risk agents sandboxed | Only 30% | | Monitoring / logging agent activity | 47% | | Enforced scoped runtime permissions | 49% | | Vendor-native security stacks dominant | 51% (OpenAI guardrails) | | Average satisfaction | 4.2 / 5 |
A more striking sub-finding: the larger the enterprise, the more likely an incident. Incident rates: 49% for 100–1,000-person companies vs. 63% for 1,000+ — yet sandbox usage among 1,000+ employee companies fell from 35% to 20%.
Money is moving too. Palo Alto Networks + CrowdStrike + Cisco collectively spent $22B on AI agent security within 12 months (Palo Alto acquiring CyberArk for $21.1B, CrowdStrike acquiring SGNL for $740M, Cisco acquiring Astrix Security for $400M).
Survey 2: The Agent Evaluation Gap (157 enterprises)
Sample: 157 enterprises with 100+ employees, single June 2026 survey.
| Key metric | Value | |---|---| | Agents deployed in past 12 months that passed internal evaluation but failed customers in production | 50% (24% multiple times) | | Fully trust automated evaluations | Only 5% | | Biggest weakness: "evaluation–real-world outcome misalignment" | 29% | | Allowed or engineering "zero-human automated deployment" | 66% (34% already + 32% in progress) | | Performing real-time quality checks on actual agent output | 23% | | Primary tools | Vendor-native evals and "no dedicated tool" tied at 17% | | Large enterprises (2,500+) moving fastest to zero-human deployment | 70% |
Worse still: among enterprises where incidents actually occurred, large enterprises (2,500+) had a 54% rate vs. 48% for SMBs — the biggest players move most aggressively and still fail to catch failures.
Vendor Positions
The M&A pace shows the major players have made their bet:
- Palo Alto Networks completed its $21.1B acquisition of CyberArk on February 11 (announced last July at $25B, the largest deal in company history).
- CrowdStrike acquired runtime authorization platform SGNL for $740M and launched its first integrated product, Continuous Identity for AI Agents, on June 15.
- Cisco announced the ~$400M acquisition of non-human identity specialist Astrix Security on May 4.
- Both surveys are single June 2026 cross-sections with self-selected, mid-market-skewed samples — best read as the perspective of enterprises actively building agent security, not the largest operators.
- "5% fully trust automated evaluations" doesn't mean 5% trust all evaluation — some may rely on human review as a backstop; VentureBeat doesn't break this out.
- "66% automated deployment" doesn't mean truly zero-human — "low-risk agent" definitions may be loose.
- VentureBeat is a media outlet; samples come from Pulse Research's own channels — keep the "self-selected" qualifier in mind.
- VentureBeat — The agent security gap (n=107, 2026-07-16): https://venturebeat.com/ai/the-agent-security-gap-54-of-enterprises-have-already-had-an-ai-agent-incident-and-most-still-let-agents-share-credentials
- VentureBeat — The agent evaluation gap (n=157, 2026-07-16): https://venturebeat.com/ai/the-agent-evaluation-gap-enterprise-ai-organizations-have-a-reality-alignment-problem-not-a-coverage-problem-and-most-are-shipping-to-production-anyway
- Secondary analysis: https://valueaddvc.com/pulse/shared-api-keys-ai-agent-security-2026 and https://saassentinel.com/2026/07/12/ai-agents-are-outpacing-enterprise-ability-to-verify-them
Analysis
Breaking the illusion: "passing evaluation ≠ ready for production"
Traditional software evals — unit, integration, end-to-end testing — work for systems with bounded inputs/outputs and reproducible errors. Agents are not that: behavior depends jointly on prompt, user context, and environment state. The same prompt can yield entirely different output combinations in different contexts. Automated evaluation can, at best, tell you the agent performed well on a given set of test cases.
VentureBeat put it bluntly: *"a passing eval is not the same as a working agent."* With 50% having shipped agents that passed evaluation but failed customers, and 29% citing evaluation–reality misalignment as the top weakness, evaluation is clearly not functioning as a quality gate.
Nobody trusts the evals, but 66% auto-deploy on them
This is the core contradiction of the evaluation gap: evaluations have become the automatic gatekeeping for agent launches — but the gate itself is uncertified. Notably, the fastest movers toward zero-human deployment are not the most technically capable firms but 2,500+ person enterprises (70%), whose internal governance teams are often least familiar with new agent failure modes.
Shared credentials are the incident amplifier
54% incidents + 69% shared credentials + 30% sandboxing form a clear cascade: once one agent is compromised, shared credentials amplify the blast radius across every workflow running with it — and 70% of enterprises have no sandbox as a backstop. Machine identities already reach 82:1 (machines:humans) per CyberArk research, while most IAM systems were never designed for this ratio. Cisco's Astrix announcement states agents are now "using (and abusing)" these credentials to execute work at scale.
Where the money is going
$22B in M&A is wildly disproportionate to actual agent-security market revenue. This is a bet that agent security becomes the next IAM battleground of the cloud era. For startups, the thesis is clear: build products in agent identity, agent scoped credentials, and agent sandboxing — not another agent wrapper.
Why It Matters
1. "Passing evaluation" is no longer a quality gate for agents. Both pre-deployment evaluation and real-time production verification must be in place. 2. Shared credentials are the easiest hole to poke. If your AI agents share the company API key, review your IAM policies now. 3. Agent security is the next cloud security battleground. With $21.1B + $740M + $400M in acquisitions, incumbents have committed; the startup window is 12–24 months. 4. The "evaluation gap" is the keyword for H2 2026 AI governance, as enterprise rollouts of ChatGPT Work and Claude Code intensify.