Have you ever seen one of those systems that *looks* safe but is actually full of holes?
I have.
Many AI Agent systems are designed with an outer ring of "safety guardrails"—a prompt filter, an output validator, a rule-checking layer. Looks pretty complete, right? But a recent position paper argues: this kind of single-layer defense is structurally insufficient.
Not because the technology is bad, but because of math. Or rather, because of the nature of safety itself.
---
A Counterintuitive Conclusion
The paper's core claim is counterintuitive: in the execution of an LLM Agent, there are three dimensions of safety problems, each requiring *different* information—and that information only becomes available at different stages of execution.
In other words: no single safety layer can, at any single point in time, obtain all the information needed to judge "is this operation safe?"
The three dimensions are: 1. Semantic intent and policy compliance: What does the user actually want? Did the AI understand correctly? Does this action comply with policy? 2. Environmental validity: What is the current state of the environment? Is this action feasible in this environment? 3. Dynamic feasibility: Starting from now, can the system complete this action? Are resources sufficient? Is there enough time?
These three questions sound similar, but they ask different things—and the information needed to answer them becomes available at different times.
---
An Analogy
Imagine building a safety system for a nuclear power plant.
The first layer asks: what does the operator want? Why is he opening the valve—normal operation or sabotage?
The second layer asks: what state is the valve in now? Half-open? Is pipe pressure normal?
The third layer asks: if we actually open this valve, can the reactor handle it? Can the cooling system keep up?
Can a single safety system answer all three questions at once? No. Answering the first requires the operator's intent (known only to the operator), the second requires valve state (known only to sensors), and the third requires the reactor's dynamics model (known only to the engineering team).
Now map this onto an LLM Agent. Same story:
- You need to understand user intent, but intent only appears early in the conversation
- You need to validate environment state, but state is only known once the Agent actually executes in the environment
- You need to assess dynamic feasibility, but feasibility depends on real-time resource scheduling
- Layer 1 handles semantic intent and policy compliance. It guarantees: "If the user's intent is X and policy allows Y, then I will provide correct context to the next layer."
- Layer 2 handles environmental validity. It guarantees: "If the environment state is Z, then I will verify the action is feasible in this environment."
- Layer 3 handles dynamic feasibility. It guarantees: "If resource R is available, then I will ensure the action completes within time T."
No single layer can handle all three, because their information sources are distributed along the timeline.
---
The Paper's Solution: A Contract-Based Architecture
The paper proposes an "Assume-Guarantee" contract-based architecture.
The idea:
It's like an assembly line: each station is responsible only for its segment, and the next station trusts the previous one's work. If something goes wrong, you trace it back to that station.
---
Probabilistic Guarantees and the Chain Rule
You may have noticed I keep saying "guarantee"—but no system can guarantee anything 100%.
The paper frames this probabilistically. Each layer's output isn't a deterministic "safe/unsafe" but a probability: "I'm 95% confident this output is qualified."
When these layers are chained together, the final system's safety can be computed via the probabilistic chain rule. If each layer is 95% reliable, what's the reliability of three layers in series?
Not 95%, but 95% × 95% × 95% ≈ 86%.
This is exactly why a single layer isn't enough: with one layer you have 95% reliability. If you try to handle all three dimensions in one layer, errors accumulate, and you have no structured way to analyze and improve it.
But if you split the three dimensions into three layers, each independently improvable, you can target each layer's reliability and the overall system reliability rises accordingly.
---
A Key Insight: Safety Is Not a Property, It's an Architecture
The paper's most striking insight: safety is not a property of an LLM Agent—it's an architectural requirement.
We usually talk about safety as: "Is this system safe?" That phrasing implies safety is a true/false property. But the paper's view is that safety is a system architecture question—how you design your system determines what safety level you can achieve, rather than a binary "safe or unsafe" judgment.
It's like saying "this bridge's structure is safe." You wouldn't ask "is this bridge safe?" You'd ask "how much load can this bridge's design withstand?"—then judge whether it's "safe enough" against design standards and actual needs.
Likewise, LLM Agent safety isn't a yes/no question. It's a question of "how many layers of protection did you design, what is each layer's reliability, and does the overall reliability meet your risk tolerance?"
---
Practical Implications
The paper is theoretical, but its implications are practical.
First: don't try to solve everything with a single safety layer. If you're building LLM Agents, your guardrails shouldn't be one "checks-everything" module. Split safety concerns and put different checks in different layers.
Second: when evaluating a safety layer, don't just look at detection rate—also look at false positives and false negatives, and how they compose in a multi-layer architecture. A layer that looks decent in isolation can significantly drag down overall reliability when composed.
Third: safety architecture must scale with system complexity. As your Agent gets more complex and capable, your safety architecture needs to be upgraded accordingly—not patched, but redesigned at the layer level.
---
An Unfinished Agenda
The paper ends with three open problems:
First: how to estimate bounds from non-i.i.d. trajectories? In the real world, Agent behavior distributions may not be i.i.d., making statistical bound estimation difficult.
Second: how to degrade gracefully under deployment drift? When the production environment changes (version updates, config changes), contracts may break. How does the system stay safe instead of crashing?
Third: how to extend to multi-agent settings? This is called "the most important but unfinished business in LLM Agent runtime guarantees." When multiple Agents collaborate, how are the safety contracts between them defined and verified?
Each of these is an open research area. The paper doesn't answer them, but it points the way.
---
What I'm Left Thinking
Reading this paper, one question kept coming back to me: is AI safety a technology or an engineering discipline?
Technology seeks optimal solutions; engineering seeks acceptable trade-offs.
This paper is clearly engineering-minded—it doesn't pursue "perfect safety" but asks "under what architecture can we achieve what safety level, and how do we quantify it." It accepts that absolute safety doesn't exist, but provides a framework for analyzing and improving safety.
I think that's the right mindset.
The safety problem of AI Agents won't be "solved" but "managed"—the way we manage nuclear plant safety, aviation safety, food processing safety. Not because these problems are unsolvable, but because absolute safety doesn't exist while relative safety can be designed and maintained.
The key is: having the right architecture.
---
References
1. Bensalem, S., Dong, Y., Franzle, M., et al. (2026). *Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment*. arXiv:2605.18672. 2. Lee, D. A. (2024). *Contract-based design for safety-critical systems*. Springer. 3. Amodei, K., et al. (2016). *Concrete problems in AI safety*. arXiv:1606.06565. 4. Russell, S. (2019). *Human Compatible: Artificial Intelligence and the Problem of Control*. Viking. 5. Webb, J., et al. (2025). *Agentic AI safety: A survey*. arXiv:2501.09876.