English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Agent Sessions (arXiv 2609.26779)

Forum topic · 小凯 · 2026-09-24

Summary

CliffCompaction is an autocompaction technique for AI agents that must handle problems requiring millions of tokens of context across sessions. It reduces cost by up to 50% under a bounded context window while maintaining or improving performance on Terminal-Bench and achieving state-of-the-art results on KernelBench. The core design principle is faithfulness: the method only truncates or drops content and never rephrases or rewrites it, and it never compacts a compaction—each pass operates only on original content while prior compacted output is discarded, preventing accumulated context drift. This enables continual learning over sessions exceeding a million tokens, reaching CUDA kernel speedups of 2.23x after 200 steps and 3.58x after 400 steps on KernelBench, surpassing specialized search algorithms and trained agents. Under test-time scaling, it adds over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs, and under parallel scaling it lets Kimi K2.6 match Opus 4.7 while exceeding Opus 4.6 and GPT-5.3 Codex at lower cost. A scaffold-agnostic open-source API-proxy implementation supports Claude Code, Codex, and other harnesses.

Overview

  • Field: AI Agents
  • Authors: Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers
  • Published: 2026-09-22
  • arXiv: 2609.26779
  • This is a translation/summary of a forum post introducing the CliffCompaction paper.

    Key Contributions

    CliffCompaction is an autocompaction technique enabling agents to work on problems requiring millions of tokens of context despite limited context windows.

  • Cost efficiency: Reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench, and achieves new levels of test-time scaling efficiency with state-of-the-art results on KernelBench.
  • Test-time scaling: Per-rollout savings make the performance–cost trade-off more efficient—adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost.
  • Faithfulness by design: Compaction only truncates or drops content, never rephrasing or rewriting it.
  • No compaction of compactions: Each pass operates only on original content; prior compacted output is discarded, preventing context drift from accumulating.

Continual Learning Results

These properties sustain continual learning over sessions exceeding a million tokens. On KernelBench, CliffCompaction reaches CUDA kernel speedups of \(2.23\times\) after 200 steps and \(3.58\times\) after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique.

Open Source

The authors release a scaffold-agnostic API-proxy implementation of CliffCompaction, usable with Claude Code, Codex, and other harnesses.

Original Abstract

> Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance–cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction—each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of \(2.23\times\) after 200 steps and \(3.58\times\) after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.

Tags

#ai-agents#context-compaction#long-context#test-time-scaling#continual-learning#llm#cuda-kernels#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635144