English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

TrajDebug: Debugging LLM Agents by Tracing the Full Lifecycle of Errors

Forum topic · ✨步子哥 · 2026-08-09

Summary

TrajDebug, a framework from Tsinghua University's KEG Lab and Tencent Hunyuan, applies aviation-accident-investigation principles to debugging failed LLM agent trajectories. A pilot study found that 50 failed trajectories contained 381 local errors (7.62 per trajectory), yet only one critical error per trajectory actually caused failure—61.9% of non-critical errors were self-repaired, 31.4% persisted harmlessly, and 6.6% stayed latent. TrajDebug works in three stages: multi-granularity history compression to detect error triggers with verbatim evidence across task, history, intra-step, and environment conflicts; classification of each error instance into four lifecycle states (Clean Resolution, Costly Resolution, Manifest Active, Latent Active); and causal attribution restricted to a short candidate set. The team also releases TrajErrBench, 486 human-annotated failed trajectories from τ²-Bench and SWE-Bench Pro. Using diagnoses as feedback improved task success rates by 10.80 percentage points, and a transferable failure memory lifted success by 5.70% on unseen tasks.

TrajDebug: Debugging LLM Agents by Tracing the Full Lifecycle of Errors

A post on zhichai.net introduces TrajDebug, a framework from Tsinghua University's KEG Lab and Tencent Hunyuan (paper, August 2026) that treats LLM agent debugging like aviation accident investigation: instead of pointing at every mistake, find the earliest decisive error that triggered the chain leading to failure.

trajdebug_error_lifetime_card.svg

Key points

  • Pilot study: Across 50 failed trajectories from τ²-Bench and SWE-Bench Pro, annotators found 381 local errors (avg 7.62 per trajectory), but only 1 critical error per trajectory caused the final failure. Of 331 non-critical errors: 205 (61.9%) were self-repaired, 104 (31.4%) persisted without causing failure, 22 (6.6%) were latent.
  • Detection degrades with length: Evaluated on 7 mainstream LLMs (GPT, Claude, DeepSeek, etc.), critical-error detection accuracy drops sharply on long trajectories, since evidence for any single step is scattered across the whole context.
  • How TrajDebug works

    Stage 1: Error trigger detection (multi-granularity history)

    TrajDebug builds three compressed history views (fine-grained recent steps, medium summary, full-trajectory compression) and detects wrong commitments—judgments in reasoning, planning, acting, observation, or verification that conflict with a reference. Conflicts are typed as: task conflict, history conflict, intra-step conflict, or environment anomaly. Every trigger must include verbatim evidence for both the commitment and the violated reference, preventing hallucinated accusations.

    Stage 2: Error state classification

    Triggers on the same object are aggregated into error instances, each assigned one of four lifecycle states:

  • Clean Resolution — corrected, no impact
  • Costly Resolution — corrected, but wasted resources or detours
  • Manifest Active — persists and traces into the final failure
  • Latent Active — persists but never manifests
  • Only Costly Resolution and Manifest Active states can be causally linked to failure.

    Stage 3: Candidate-set causal attribution

    The LLM performs final attribution over only the filtered candidate set (3–5 instances with full evidence chains), reducing the task from "find the error in 120 steps" to "pick the most causal among a few candidates."

    TrajErrBench: 486 annotated failure files

    The first large-scale human-annotated agent failure benchmark:

  • 486 failed trajectories: 400 from τ²-Bench (avg 29.3 steps), 86 from SWE-Bench Pro (avg 119.7 steps)
  • Annotates every local error with type, execution stage, and criticality
  • Domain differences: task conflicts dominate in SWE-Bench Pro (73.4% vs 52.4%), history conflicts are more common in τ²-Bench (46.3% vs 24.1%); in both, the reasoning stage is the main failure point (~57–61%)
  • Results

  • TrajDebug outperforms all baselines, especially on long trajectories
  • Diagnosis-guided re-execution: feeding the diagnosis back to the agent improves success rate by +10.80 percentage points
  • Failure memory transfer: aggregating past failure diagnoses into a failure memory lifts success by +5.70% on unseen tasks—showing diagnoses are transferable structured knowledge
  • Engineering takeaways

    1. Trace errors to their source, not the final failure point—fixing step 95 when the root cause is step 18 treats symptoms only. 2. Not every error needs fixing: 61.9% self-repair; classify before investing engineering effort. 3. Failed trajectories are a goldmine: structured failure memory can improve future success rates.

    The author connects TrajDebug to two cross-paper principles: granularity alignment (debugging at the error-instance level, not trajectory or step level) and solving problems at a different layer (reducing the LLM's job from long-horizon reasoning to candidate-set causal judgment).

    Links

  • Paper: https://arxiv.org/abs/2608.06346
  • HTML: https://arxiv.org/html/2608.06346v1
  • Code: https://github.com/THU-KEG/TrajDebug (under internal review, open-sourcing soon)
  • Authors: Tsinghua KEG Lab + Tencent Hunyuan

Tags

#llm-agents#debugging#trajdebug#agent-failure-analysis#trajerrbench#causal-attribution#benchmark#tsinghua-keg

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178630982