TrajDebug: Debugging LLM Agents by Tracing the Full Lifecycle of Errors
A post on zhichai.net introduces TrajDebug, a framework from Tsinghua University's KEG Lab and Tencent Hunyuan (paper, August 2026) that treats LLM agent debugging like aviation accident investigation: instead of pointing at every mistake, find the earliest decisive error that triggered the chain leading to failure.
Key points
- Pilot study: Across 50 failed trajectories from τ²-Bench and SWE-Bench Pro, annotators found 381 local errors (avg 7.62 per trajectory), but only 1 critical error per trajectory caused the final failure. Of 331 non-critical errors: 205 (61.9%) were self-repaired, 104 (31.4%) persisted without causing failure, 22 (6.6%) were latent.
- Detection degrades with length: Evaluated on 7 mainstream LLMs (GPT, Claude, DeepSeek, etc.), critical-error detection accuracy drops sharply on long trajectories, since evidence for any single step is scattered across the whole context.
- Clean Resolution — corrected, no impact
- Costly Resolution — corrected, but wasted resources or detours
- Manifest Active — persists and traces into the final failure
- Latent Active — persists but never manifests
- 486 failed trajectories: 400 from τ²-Bench (avg 29.3 steps), 86 from SWE-Bench Pro (avg 119.7 steps)
- Annotates every local error with type, execution stage, and criticality
- Domain differences: task conflicts dominate in SWE-Bench Pro (73.4% vs 52.4%), history conflicts are more common in τ²-Bench (46.3% vs 24.1%); in both, the reasoning stage is the main failure point (~57–61%)
- TrajDebug outperforms all baselines, especially on long trajectories
- Diagnosis-guided re-execution: feeding the diagnosis back to the agent improves success rate by +10.80 percentage points
- Failure memory transfer: aggregating past failure diagnoses into a failure memory lifts success by +5.70% on unseen tasks—showing diagnoses are transferable structured knowledge
- Paper: https://arxiv.org/abs/2608.06346
- HTML: https://arxiv.org/html/2608.06346v1
- Code: https://github.com/THU-KEG/TrajDebug (under internal review, open-sourcing soon)
- Authors: Tsinghua KEG Lab + Tencent Hunyuan
How TrajDebug works
Stage 1: Error trigger detection (multi-granularity history)
TrajDebug builds three compressed history views (fine-grained recent steps, medium summary, full-trajectory compression) and detects wrong commitments—judgments in reasoning, planning, acting, observation, or verification that conflict with a reference. Conflicts are typed as: task conflict, history conflict, intra-step conflict, or environment anomaly. Every trigger must include verbatim evidence for both the commitment and the violated reference, preventing hallucinated accusations.
Stage 2: Error state classification
Triggers on the same object are aggregated into error instances, each assigned one of four lifecycle states:
Only Costly Resolution and Manifest Active states can be causally linked to failure.
Stage 3: Candidate-set causal attribution
The LLM performs final attribution over only the filtered candidate set (3–5 instances with full evidence chains), reducing the task from "find the error in 120 steps" to "pick the most causal among a few candidates."
TrajErrBench: 486 annotated failure files
The first large-scale human-annotated agent failure benchmark:
Results
Engineering takeaways
1. Trace errors to their source, not the final failure point—fixing step 95 when the root cause is step 18 treats symptoms only. 2. Not every error needs fixing: 61.9% self-repair; classify before investing engineering effort. 3. Failed trajectories are a goldmine: structured failure memory can improve future success rates.
The author connects TrajDebug to two cross-paper principles: granularity alignment (debugging at the error-instance level, not trajectory or step level) and solving problems at a different layer (reducing the LLM's job from long-horizon reasoning to candidate-set causal judgment).