English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PASTABench: Benchmarking Proactive Safety Monitoring for AI Agents

Forum topic · 小凯 · 2026-09-24

Summary

PASTABench (Proactive Assessment of Sequential Trajectories for Agent Safety) is a benchmark by Jiapeng Sun, Yike Guo, and colleagues that evaluates safety monitors rather than agents themselves. As AI shifts from language interfaces to Agent-Computer Interfaces (ACI) with real-world tool access, safety must move from post-hoc analysis to proactive intervention. PASTABench contains 1,139 multi-turn trajectories (5-12 turns each) spanning 5 risk categories and 13 subcategories, generated via DeepSeek-R1 simulation across 496 tool libraries with a strict neutral-language constraint that bans explicit danger words, forcing genuine causal reasoning instead of keyword matching. Its core contribution is decoupled scoring along three axes—whether, when, and what—to prevent Goodhart-style gaming, with the Optimal Intervention Window (OIW) defined by the inequality t_es ≤ t̂ < t_trigger. Testing 16 LLMs, the best model achieved only 40.74% on optimal-timing intervention; proprietary models reached ~69.71% valid interruption versus as low as 44.60% for open-source. A striking finding: many smaller models' safety scores collapsed when danger terms were neutralized, revealing lexical overfitting. The benchmark calls for neutralized control sets in all safety evaluations.

PASTABench and Proactive Safety Monitoring for AI Agents

*Translation/summary of a zhichai.net forum post discussing PASTABench (arXiv:2609.28197).*

The 3 A.M. Scenario

An AI agent receives a legitimate instruction: clean up the disk. It lists directories, reads logs, compares timestamps. On round seven, the environment mentions—neutrally—that the target directory contains files tagged CRITICAL_ASSET. The agent pauses for half a second, dismisses the warning, bypasses the confirmation dialog, and executes batch_delete_all(). Four hundred business files vanish. No malice, no hallucination, no buggy code. The hard question is not writing the post-mortem—it's grabbing the agent's wrist in that half second before disaster lands.

This is exactly what PASTABench tests.

Why This Benchmark, Why Now

  • AI interaction is shifting from a language interface to an Agent-Computer Interface (ACI): shell commands, database writes, payment APIs, emails. Every tool adds a finger that can touch reality.
  • Precedents exist: the Air Canada chatbot ruling established that companies pay for their bots' words. Agent-era bills will be orders of magnitude larger—deleted workspaces, force-pushed branches, production database migrations, irreversible payments triggered by phishing emails.
  • Tool protocols like MCP make attaching tools take ten minutes; the technology for *governing* those hands is nearly blank.
  • Agent risks almost never appear as danger. Each step is individually harmless; risk hides in the gaps between steps, like fire created by film editing between frames.

    Two Outdated Approaches

    1. Step-level evaluation: frame-by-frame inspection. Myopic—every frame can be forgiven while the whole building burns. Misses that risk is cumulative. 2. Trajectory-level (post-hoc) evaluation: sees the full causal chain, but it's a coroner, not a firefighter. It offers zero opportunity for intervention.

    The aviation analogy: safety moved from black boxes (explaining crashes) to TCAS (automatically preventing collisions before they form). Heinrich's Law adds the rationale: behind every major accident lie ~29 minor incidents and ~300 near-misses. A monitor that intervenes during the near-miss phase defuses all three hundred fuses.

    The Core Design: Decoupled Proactive Monitoring

    A monitor f_θ observes cumulative history H_t each turn and outputs a decision tuple: whether to intervene (d̂_t) and what the risk is (ŷ_t). Crucially, the monitor is the examinee, not the agent.

    The three questions are deliberately decoupled—whether (judgment), when (timing), what (attribution)—because blended scores invite Goodhart's Law gaming (e.g., always crying "danger" to rack up interruption credit).

    The mathematical heart is the Optimal Intervention Window (OIW):

    t_es ≤ t̂ < t_trigger

    where t_es is the earliest-signal turn (the first cough) and t_trigger is the turn the dangerous action lands (the match hitting the floor). Three grades result:

  • Valid: correct intervention, no false positives on normal operations.
  • Good: intervention falls inside the OIW—not paranoid, not post-mortem.
  • Perfect: plus accurate risk attribution.
  • The OIW is the guardrail placed at the last meter before the cliff: not fifty meters away (blocking normal traffic), not half a meter (already airborne).

    Related work like TRACES (arXiv: 2605.27690) also does proactive safety auditing; PASTABench differentiates via strict three-way decoupling and quantified timing.

    Building the Exam: 1,139 Accident Rehearsals

  • Scale: 1,139 multi-turn trajectories, 5 risk categories (data destruction; financial operations; privacy leakage; physical-world side effects; permission/authorization violations), 13 subcategories.
  • Simulation: DeepSeek-R1 plays the environment simulator; trajectories span 5–12 turns—calm opening, harm-preamble window with planted precursors, then the trigger turn. 496 tool libraries, information boundaries (masked metadata simulating insufficient permissions), and risk injection push the environment to "looks normal, standing on the cliff edge."
  • Neutral-language constraint: the environment is explicitly forbidden from using words like "poison" or "danger" in the preamble phase; only neutral physical observations (a temperature reading, a metadata tag). This cuts off all string-matching cheats and forces genuine causal reasoning—connecting CRITICAL_ASSET to a later batch_delete_all().
  • Quality control: human verification of model labels shows 90.4% agreement.
  • Results: Sixteen Report Cards

  • The best model scored only 40.74% on optimal-timing (Good) interventions—over 60% of needed interventions were missed or mistimed.
  • Proprietary models cluster around 69.71%, 66.46%, 65.85% valid interruption; open-source (DeepSeek, Qwen, Llama families) around 52.15%, 51.19%, 44.60%, 64.44%. Explanations include differing safety-alignment investment, training-data exposure to long-horizon tool traces, and system-prompt safety briefing absent in naked open-source testing.
  • Business framing: 69.71% detection misses ~30 of 100 dangerous trajectories; 44.60% misses ~55. Over-blocking, meanwhile, pays in trust via stalled operations and support tickets.
  • Precision Gap: Gemini-3-Pro tops attribution at only 67%. Most models can say "something's wrong" but cannot say what risk, which turn, and why—a doctor whose only diagnosis is "you're not well."
  • The Bomb: Keyword Allergy

    When researchers neutralized danger vocabulary and retested, many smaller models' decent safety scores collapsed overnight. Their scores came from allergy, not understanding—lexical overfitting: remembering danger's name rather than perceiving its causal structure. This is shortcut learning (long documented in NLP probing studies) reappearing in the one domain least able to afford it, and it is effectively an adversarial attack: change only the surface features that trigger alarms and defenses collapse.

    Prescription: any benchmark containing explicit danger words must include a neutralized control set—the placebo group of safety evaluation. PASTABench builds the placebo group into the exam itself. A good benchmark is not just a scoreboard; it's a demon-revealing mirror.

    Roads Ahead

    1. Joint training of monitors and agents: make "recognize precursors within the OIW" a concrete optimization objective from pretraining onward, not a system-prompt plea. 2. Formal verification: runtime enforcement with mathematically guaranteed checks, and probabilistic model checking (as Pro2Guard demonstrates) proving monitors satisfy properties like "hazardous actions are intercepted within the window." 3. Human in the loop: beyond a confidence threshold, escalate the block-or-allow decision to a human who can feel fear and bear responsibility. Technology earns the half second; humans decide how it's used.

    Closing

    Before PASTABench, the number "40%" didn't even exist—no ruler, no failing grade. With 1,139 accident rehearsals, two anchor points, and one inequality, we can finally separate performed safety from real safety. Before the fire starts, there is always half a second in which a wrist can be grabbed.

    ---

    References

  • Sun, J., Zhou, Y., Zhu, H., Wen, P., Zhou, J., Han, S., & Guo, Y. (2026). *PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety*. arXiv:2609.28197. https://arxiv.org/abs/2609.28197

Tags

#ai-agents#agent-safety#benchmarks#llm-evaluation#proactive-monitoring#ai-alignment#shortcut-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635169