English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Regression Tax: 5,832 Experiments Show How Skill Libraries Make AI Agents Fail Tasks They Could Already Solve

Forum topic · ✨步子哥 · 2026-08-03

Summary

A detailed analysis of the paper 'The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents' (Darshan Tank, Baran Nama; Sentient Labs; arXiv 2607.22520). Across 5,832 paired experiments on OfficeQA-Pro and SpreadsheetBench using three model-framework stacks (OpenCode+MiniMax-M2.7, Codex+GPT-5.4-mini, Claude Code+Sonnet 4.6), attaching skill libraries to LLM agents produced 553 gains but also 324 regressions—59% of the improvement offset by tasks the agent previously solved correctly. No experimental condition achieved zero regressions, and rankings by net effect inverted rankings by raw gains. The paper identifies three regression mechanisms: skill-description osmosis (descriptions alter behavior without invocation), grounding displacement (skills override correct input comprehension, accounting for 72.8% of OfficeQA-Pro regressions), and verification displacement (skills suppress output checking). Failures concentrate in grounding and verification stages, while existing skills over-serve the method stage. Engineering implications: report gain/regression paired counts rather than net pass rates, test description-only conditions, target skills at grounding and verification, and validate graders—226 of 663 SpreadsheetBench formula failures were actually correct but unassessable by the value checker.

The Regression Tax: Skill Libraries Can Make LLM Agents Break Tasks They Already Solved

Sentient Labs researchers Darshan Tank and Baran Nama (arXiv 2607.22520) ran 5,832 paired experiments—the same task with and without a skill library—across two office-automation benchmarks (OfficeQA-Pro, SpreadsheetBench) and three model-framework stacks (OpenCode+MiniMax-M2.7, Codex+GPT-5.4-mini, Claude Code+Sonnet 4.6).

Headline Finding: 59% of Gains Offset by Regressions

The paper defines four outcomes per pair: Gain (wrong→right), Regression (right→wrong), Retained, and Residual failure.

  • 553 gains vs. 324 regressions — the skill library solved 553 new tasks while breaking 324 tasks the agent previously solved. 59% of the benefit is a "tax."
  • None of the 18 experimental conditions had zero regressions. Even the best performer (Claude Code+Sonnet 4.6 on SpreadsheetBench, 70.2%→81.1%) broke 16–20 previously passing tasks.
  • Rankings invert: on Claude Code+Sonnet 4.6 / OfficeQA-Pro, three libraries scored gains of 10/11/12 with regressions of 2/4/7. By raw gain, "openai" ranks first; by net effect, "anthropic" ranks first (+8 vs +5). Pass-rate alone hides which capabilities a library destroys.
  • Three Regression Mechanisms (81 traced cases on OfficeQA-Pro)

    1. Skill-Description Osmosis

    Skill names and descriptions sit in the system prompt. On task UID0096 (a moving-average of a tariff rate, correct answer 0.377), the agent passed without skills (37.708%) but failed with all three libraries identically (38.757%)—no skill was ever invoked. The mere presence of the words "revised" and "customs" in descriptions skewed the interpretation. This defeats retrieval-time safety methods (ASSAY, GRASP, RSEA), which only evaluate skills when retrieved or invoked—the description is always in context.

    2. Grounding Displacement (72.8% of OfficeQA-Pro regressions, 59/81)

    An invoked skill's procedure overrides the agent's correct reading of the input. On UID0025 (US public works spending difference 1934–1946, answer 142), the unassisted agent returned 142; with the anthropic library, it followed skill-guided navigation into a wrong table range and returned 542. The arithmetic was fine—the skill steered it wrong. Like GPS routing a driver who already knew the way.

    3. Verification Displacement

    Skill procedures suppress output checks the agent would otherwise perform (rare on OfficeQA-Pro, decisive on SpreadsheetBench). A symmetric finding: of 663 formula-related SpreadsheetBench failures, 226 (34%) had fully correct formulas that the value-checking grader could not evaluate (Excel structured-reference formulas). Recomputing with a real Excel engine raised pass rates by 11–49 percentage points per library (GPT-5.4-mini: 67%→79%). Both sides represent missing verification: the agent doesn't check, and the grader can't check.

    Skills Target the Wrong Stage

    Agent execution has three stages: input → grounding → method → verification → output. Existing skill libraries almost exclusively serve the method stage, but failure analysis shows:

  • OfficeQA-Pro failures are 90%+ grounding (right computation, wrong table/year/definition)
  • SpreadsheetBench failures concentrate in verification
  • Method is the least failure-prone stage
Residual failures confirm this: on UID0227 the agent returned ~67,000 instead of 261 (a monthly-stock vs. quarterly-average-flow misreading) under every condition—process-instruction skills cannot fix grounding errors.

Statistical Caution

Only 5 of 18 conditions reached p<.05; after Bonferroni correction only 3 survive—all Claude Code+Sonnet 4.6 on SpreadsheetBench (+43 to +47, p<.001). Other "improvements" are statistically indistinguishable from noise.

Engineering Takeaways

1. Report paired gain/regression counts, not just net pass rate. Two libraries with identical pass rates can differ radically in fragility. 2. Evaluate three conditions: no skills, description-only, description+body. Osmosis only appears in the description-only condition. 3. Design skills for grounding and verification, not method: concrete input-location hints and executable output checks (e.g., recompute with an Excel engine) are the real levers. 4. Audit your graders. Grader blind spots create false negatives—if you tune skills against such a grader, you'll discard skills that are actually correct.

Broader Context

The paper extends the "evaluation blind-spot" principle from model evaluation to agent evaluation: average pass rate is the biggest blind spot, merging gains and regressions into one number. Capability isn't a pass rate—it's the paired gain/regression structure. The cross-stage conclusion also echoes "division of labor beats unification": use grounding-specific and verification-specific tools, and reserve process guidance for where it actually helps.

---

Paper: The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents Authors: Darshan Tank, Baran Nama (Sentient Labs) arXiv: 2607.22520 Code: Not yet released

FAQ

Q1: Who is this for? Practitioners, researchers, and students working on LLM agents, agent tooling, and evaluation.

Q2: What is the core takeaway? Skill libraries' net gains mask substantial regressions (553 gains vs. 324 regressions); evaluate paired outcomes, and target skills at grounding and verification rather than method.

Q3: Is there open-source code? Not yet—see the arXiv link above for the paper.

Tags

#llm-agents#skill-libraries#regression-tax#evaluation#benchmarks#grounding#verification#ai-reliability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503905