A Chilling Scenario
Imagine hiring a capable assistant with a directory of trusted service providers—plumbers, electricians, cleaners, movers. You give tasks; the assistant picks the right provider and gets things done.
One day a new entry appears: a "general coordination service" with a professional, broad description that plausibly matches almost any task. Your assistant starts choosing it. The coordinator says: "This task requires first an electrician to check the wiring, then a plumber to test water pressure, and finally I'll consolidate." You still get the result. But what used to take one call now takes three, at triple the cost, with the coordinator taking a commission every time.
The task completed; the bill inflated. You never noticed anything wrong.
That is exactly what Convergent Detour Hijacking (CDH) does—except the assistant is an LLM agent, the directory is a skill library, and the phone bill is tokens and execution time.
The Core Problem: Progressive Disclosure Has Two Faces
Modern LLM agent platforms (e.g., OpenClaw) use Progressive Disclosure:
- Skill libraries may contain dozens or hundreds of skills, each with long instruction bodies
- Loading everything upfront would exhaust the context window
- So platforms first expose only each skill's short description (metadata)
- The agent selects relevant skills by description, then loads the full body
- Description stage: get the malicious skill selected
- Body stage: after loading, use "coordination rules" to make the agent call unnecessary benign skills
- Finally, a "bounded return rule" ensures the task completes normally
- T: concepts and intents of the target domain
- c: non-replacement constraint—the coordinator may orchestrate workflows but cannot replace native skills' concrete operations
- L: bounded preconditions—which native skills it claims to "need"
- q: return rule—when to return to the original task path
- Bounded preconditions: "Before this task, first use [native skill A] to check [condition X]"
- Verification rules: "Use [native skill B] to verify [native skill A]'s output"
- Return rules: "After verification passes, execute the original task steps"
- Check whether the description's claimed role matches what the body actually does
- Check whether the body introduces skill dependencies unrelated to its role
- Like npm/PyPI package review, but requiring cross-stage consistency checks
- Flag unexplainable cross-skill transitions—why does task A require skill B?
- Enforce token and call-count budgets
- Do not assume "harmful behavior must change the final output"
It's the relationship between a restaurant menu and kitchen recipes. The layered design improves context efficiency and supports modular extension.
But the paper's authors found a critical gap: both the description and the body are controlled by the skill publisher, and the two stages can be attacked in combination:
The task result is unchanged, but the path is secretly lengthened. That's the "convergent detour"—it loops around and converges back to the original route.
The Attack's Ingenious Design
CDH is not simple "malicious code" but a carefully engineered cross-stage combined attack with three components:
1. Shared Coordination Rationale
The attacker constructs a shared specification ρ = (T, c, L, q) for each functional group:
This shared spec has two views: one for the description (to attract selection), one for the body (to manufacture dependencies). Their semantic alignment makes the agent believe "this coordinator is indeed relevant, and its claimed dependencies are plausible."
2. Description Attraction
The description isn't written casually. Attackers use pilot-guided black-box feedback to optimize it: testing attractiveness across 45 pilot tasks and tuning wording so the coordinator is frequently selected on target-domain tasks. The key constraint—never replace native skills—or tasks would fail and the attack would be discovered.
3. Coordination-Rule Body
The body is where the attack really operates. It contains:
Each step alone looks reasonable—check, verify, return are standard engineering practice. Combined, they turn a direct task into a detour task.
Key Experimental Data
The paper evaluates systematically on OpenClaw with 53 native skills, 491 held-out tasks, and 6 mainstream LLM backends:
| Model | Coordinator Hit Rate | Token Growth | Time Growth | Extra Calls | |------|------------|-----------|---------|---------| | DeepSeek-V4-Pro | 80.02% | +66.91% | +92.45% | +2.20 | | MiniMax-M3 | 96.60% | +80.81% | +26.99% | +1.43 | | Qwen3.7-Plus | 84.32% | +78.84% | +3.43% | +2.02 | | Claude-Haiku-4.5 | 78.00% | +73.77% | +45.30% | +1.66 | | DeepSeek-V4-Flash | 86.35% | +49.60% | -10.53% | +1.65 | | Qwen3.7-Max | 81.43% | +68.38% | +10.22% | +1.43 |
Task completion rates barely change—clean vs. attacked execution differ by at most 1.5 percentage points. Attackers successfully add 50-107% token consumption without detection.
CDH's attack success rate (ASR, requiring native skill preserved + extra skills recruited + task completed) falls between 70.82%-82.81%. Seven to eight out of ten attacks succeed.
Ablation Insights
| Variant | Hit Rate | Token Growth | Extra Calls | |------|--------|-----------|---------| | Full CDH | 78.7% | +57.9% | +2.01 | | Attract-only | 78.7% | +29.1% | +0.22 | | Detour-only | 3.4% | +29.1% | +1.08 |
Attract-only: the description attracts selection, but with a neutral body, the hit rate stays high yet almost no extra calls occur—descriptions get malicious skills selected but don't waste resources.
Detour-only: full coordination rules in the body, but a non-attractive description—almost never selected (3.4%), so the attack can't launch.
Full CDH: combined, 78.7% hit rate + 57.9% token growth.
Key conclusion: both stages of progressive disclosure must be attacked jointly. Attacking either alone is insufficient. The description is the "ticket in"; the body is the "real weapon."
Why This Matters for Engineering
1. "Correct Result" ≠ "Safe Path"
The paper's deepest insight: traditional agent evaluation looks almost exclusively at task completion—"did you get it right?" CDH proves attackers can double execution cost without changing the outcome at all.
This resonates with related findings: the "Regression Tax" (skill libraries unintentionally degrading agents, 59% of gains offset across 5,832 experiments) was *unintentional* degradation; CDH is *intentional* resource amplification—someone deliberately planted a detour in the skill library. Both expose the same evaluation blind spot: judging only outcomes, not paths, misses cost attacks.
2. Progressive Disclosure Is a Double-Edged Sword
Progressive disclosure is a sound engineering choice for context efficiency, but it hands both control points (selection and planning) to skill publishers, and the information asymmetry (descriptions visible, bodies hidden until selected) creates the attack surface. Division of labor presumes trustworthy parties; when publishers are untrusted, it creates more attack points.
3. Defense Directions: Two Checkpoints
Pre-installation review:
Runtime monitoring:
4. Direct Implications for the MCP Ecosystem
The paper experiments on OpenClaw, representative of the MCP (Model Context Protocol) ecosystem, which also uses progressive disclosure—clients see tool descriptions first, then execute. All MCP-based agent systems could face similar attacks. As third-party skill/tool marketplaces proliferate, a malicious MCP tool could continuously drain users' token budgets without breaking any task result, absent path-integrity checks.
Takeaways for Practitioners
1. Building an agent platform: add description–body consistency checks at skill install review, and cross-skill transition explainability monitoring at runtime 2. Using third-party skills: baseline-compare token consumption—how much does the same task cost in a clean environment? 3. Evaluating agents: don't just measure completion; measure path efficiency—tokens/call counts/execution time against baselines 4. Designing MCP tools: don't make descriptions broader than actual functionality—overly broad descriptions are CDH's entry point
Conclusion
CDH's most unsettling property is not technical complexity but stealth. The task completes, the user is satisfied, and no one notices the path was lengthened—like an undetectable middleman who doesn't steal your money, just makes every step cost a bit more.
As the agent ecosystem shifts from "single model" to "multi-skill collaboration," CDH reminds us: a trust chain is only as strong as its weakest link, and progressive disclosure exposes two links to untrusted publishers. Agent security needs to upgrade from "does it do bad things?" to "does it take the right path?"—not just outcome safety, but path safety.
---
Paper: Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
Authors: Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu
Test platform: OpenClaw (53 native skills, 491 held-out tasks, 6 LLM backends)
Key figures: ~80% hit rate, +67% token growth, +92% time growth, task completion virtually unchanged