Summary
TokenPilot (LightMem2) is a two-tier context management framework that cuts LLM agent inference cost by 61-87% while matching or improving task performance, without breaking prompt prefix stability. The global tier performs ingestion-aware compaction: it partitions messages into internal intent vs open-world environment feedback, applies prefix stabilization (replacing volatile runtime markers such as working directory, timestamp, and session id with static placeholders), and lossy-but-recoverable observation reduction (HTML slimming, output truncation, image downsampling, deduplication) backed by an external artifact registry. The local tier performs lifecycle-aware eviction using a three-state model (active, completed, evictable) and a lightweight zero-shot residual-utility estimator (Qwen3.5-35B-A3B) that runs in batched turns to avoid per-turn text mutations. On PinchBench and Claw-Eval, TokenPilot is the only method that simultaneously lowers cost and improves overall score; in continuous mode it controls the cost explosion from $5.16 to $10.58 versus vanilla's $81.52. Released as an OpenClaw plugin with conservative, normal, and aggressive presets.
Key points
- Paper: TokenPilot: Cache-Efficient Context Management for LLM Agents — Xu, Xue, Chen, et al. (Zhejiang University, UESTC, Xidian, HomologyAI). arXiv: https://arxiv.org/abs/2606.17016 ; Code: https://github.com/zjunlp/LightMem2
- Problem framing: Existing compression or eviction methods (LLMLingua-2, SelectiveContext, LCM, Pichay, MemoBrain, MemOS) mutate the prompt layout and destroy the byte-level prefix, causing KV cache invalidation. The cost model is
K(C') = α·|C'_hit| + |C'_miss| with α ≪ 1, so preserving cache hit continuity matters more than raw token reduction.
- Two-tier architecture:
- *Global (Ingestion-Aware Compaction)* partitions messages into
Ω_int (task prompts, thoughts, tool calls, responses) and Ω_env (unmanaged external feedback such as HTML, logs, files). It normalizes volatile runtime markers to placeholders (<WORKDIR>, <TimeStamp>, <AGENT ID>) and relocates tool definitions to the end of the dynamic block. Then it applies lossy reduction with content-hash-indexed external artifacts for recovery.
- *Local (Lifecycle-Aware Eviction)* tracks chunks across
active → completed → evictable and uses a zero-shot residual-utility estimator (Qwen3.5-35B-A3B) in a batched-turn schedule (B=3) to decide when completed tasks can be safely evicted without triggering compensating re-reads.
- Prefix stabilization gains (cache hit rate):
- PinchBench: 38.7% → 79.2% (77.24% of tasks gain 5,120+ token cache warm-up)
- Claw-Eval: 67.2% → 83.1% (59.6% of tasks gain 6,144 token warm-up)
- Reduction controls:
triggerMinChars=2200, maxToolChars=1200, global truncation 50k chars; HTML slimming, 600-char head + 400-char tail tool output truncation, image downsampling (bitmap ≤100KB, SVG ≤50KB), 5-call dedupe cap.
- Results:
- PinchBench Isolated: Overall 81.0 (vs 80.5 vanilla), Cost $3.22 (vs $8.31, −61.3%), Cache Miss 1.933M (vs 8.753M, −77.9%).
- PinchBench Continuous: Overall 81.3 (vs 79.2), Cost $2.79 (vs $7.24, −61.5%).
- Claw-Eval Isolated: Overall 63.1 (vs 64.5), Cost $2.27 (vs $5.16, −56%).
- Claw-Eval Continuous: Overall 60.8, Cost $10.58 (vs vanilla $81.52, −87%; MemOS $24.12; LLMLingua-2 $82.91). TokenPilot is the only baseline that reduces cost while matching or exceeding vanilla performance.
- Ablations:
- Progressive: +IAC drops cost 41.7% and cache miss 73.3%; +LAE further drops cost 33.9% and cache read 68.0%.
- Removing the recovery tool: −4.7 Overall, +40.4% cost (compensating retries).
- Batch size B=3 is the validated sweet spot; B=1 truncates too early, B≥5 inflates memory and latency.
- Integration: Shipped as a configurable OpenClaw plugin (LightMem2) with
conservative, normal, and aggressive presets, runtime commands (/lightmem2 status|report|doctor|visual|mode), and planned ports to Codex CLI and Claude Code.
- Limitations: estimator misclassification risk on sparse/ambiguous interactions; τ and B may need per-deployment tuning; requires backend prefix-cache support; assumes same-category tasks share continuous sessions, which can reduce prefix reuse for heterogeneous task mixes.
- Future directions: multi-agent cross-context KV-cache topology, adaptive thresholds, distilling evicted tasks into reusable skills (linked to
memory.autoDistill and the SkillClaw direction), and larger-scale industrial evaluation beyond the current 123/161 task subsets.
Takeaway
TokenPilot reframes agent cost optimization as preserving prefix-cache continuity rather than minimizing token counts. The three durable ideas are byte-level prefix stabilization, ingestion-time denoising of open-world observations, and batched residual-utility-driven eviction.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178207996