English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

Forum topic · 小凯 · 2026-08-29

Summary

RedEvoAgent is a black-box red-teaming agent designed to stress-test LLM-based agents operating in production-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes that pose greater risks than unsafe text generation. Unlike prior automatic red-teaming methods that rely on fixed attacks or trajectory-based retrieval—which can reuse misleading experiences due to retrieval bias and unclear tool attribution—RedEvoAgent distills cross-case attack trajectories into concise, human-readable attack skills. These skills adaptively evolve through tool-effectiveness profiling and decision-level tool attribution, while a verification ratchet retains only updates that improve verified attack success. Experiments across multiple benchmarks, target models, and execution environments show that RedEvoAgent outperforms both fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution environments. Paper: arXiv 2508.11367.

Overview

  • Field: Machine Learning (ML)
  • Authors: Junjie Zhang, Hui Liu, Kecheng Chen
  • arXiv: 2508.11367
  • Abstract

    LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability.

    Key Contributions

  • RedEvoAgent: a black-box red-teaming agent that distills cross-case attack trajectories into concise, human-readable attack skills.
  • Adaptive skill evolution: attack skills evolve via tool-effectiveness profiling and decision-level tool attribution for skill updates.
  • Verification ratchet: only skill updates that improve verified performance are retained, filtering out misleading experiences.

Results

Experiments across multiple benchmarks, target models, and target execution environments demonstrate that RedEvoAgent outperforms both fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution environments.

---

*Auto-collected on 2026-08-29.*

Tags

#red-teaming#llm-agents#jailbreak#ai-safety#paper#arxiv#machine-learning#agent-security

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634193