The Regression Tax: Skill Libraries Can Make LLM Agents Break Tasks They Already Solved
Sentient Labs researchers Darshan Tank and Baran Nama (arXiv 2607.22520) ran 5,832 paired experiments—the same task with and without a skill library—across two office-automation benchmarks (OfficeQA-Pro, SpreadsheetBench) and three model-framework stacks (OpenCode+MiniMax-M2.7, Codex+GPT-5.4-mini, Claude Code+Sonnet 4.6).
Headline Finding: 59% of Gains Offset by Regressions
The paper defines four outcomes per pair: Gain (wrong→right), Regression (right→wrong), Retained, and Residual failure.
- 553 gains vs. 324 regressions — the skill library solved 553 new tasks while breaking 324 tasks the agent previously solved. 59% of the benefit is a "tax."
- None of the 18 experimental conditions had zero regressions. Even the best performer (Claude Code+Sonnet 4.6 on SpreadsheetBench, 70.2%→81.1%) broke 16–20 previously passing tasks.
- Rankings invert: on Claude Code+Sonnet 4.6 / OfficeQA-Pro, three libraries scored gains of 10/11/12 with regressions of 2/4/7. By raw gain, "openai" ranks first; by net effect, "anthropic" ranks first (+8 vs +5). Pass-rate alone hides which capabilities a library destroys.
- OfficeQA-Pro failures are 90%+ grounding (right computation, wrong table/year/definition)
- SpreadsheetBench failures concentrate in verification
- Method is the least failure-prone stage
Three Regression Mechanisms (81 traced cases on OfficeQA-Pro)
1. Skill-Description Osmosis
Skill names and descriptions sit in the system prompt. On task UID0096 (a moving-average of a tariff rate, correct answer 0.377), the agent passed without skills (37.708%) but failed with all three libraries identically (38.757%)—no skill was ever invoked. The mere presence of the words "revised" and "customs" in descriptions skewed the interpretation. This defeats retrieval-time safety methods (ASSAY, GRASP, RSEA), which only evaluate skills when retrieved or invoked—the description is always in context.2. Grounding Displacement (72.8% of OfficeQA-Pro regressions, 59/81)
An invoked skill's procedure overrides the agent's correct reading of the input. On UID0025 (US public works spending difference 1934–1946, answer 142), the unassisted agent returned 142; with the anthropic library, it followed skill-guided navigation into a wrong table range and returned 542. The arithmetic was fine—the skill steered it wrong. Like GPS routing a driver who already knew the way.3. Verification Displacement
Skill procedures suppress output checks the agent would otherwise perform (rare on OfficeQA-Pro, decisive on SpreadsheetBench). A symmetric finding: of 663 formula-related SpreadsheetBench failures, 226 (34%) had fully correct formulas that the value-checking grader could not evaluate (Excel structured-reference formulas). Recomputing with a real Excel engine raised pass rates by 11–49 percentage points per library (GPT-5.4-mini: 67%→79%). Both sides represent missing verification: the agent doesn't check, and the grader can't check.Skills Target the Wrong Stage
Agent execution has three stages: input → grounding → method → verification → output. Existing skill libraries almost exclusively serve the method stage, but failure analysis shows:
Statistical Caution
Only 5 of 18 conditions reached p<.05; after Bonferroni correction only 3 survive—all Claude Code+Sonnet 4.6 on SpreadsheetBench (+43 to +47, p<.001). Other "improvements" are statistically indistinguishable from noise.
Engineering Takeaways
1. Report paired gain/regression counts, not just net pass rate. Two libraries with identical pass rates can differ radically in fragility. 2. Evaluate three conditions: no skills, description-only, description+body. Osmosis only appears in the description-only condition. 3. Design skills for grounding and verification, not method: concrete input-location hints and executable output checks (e.g., recompute with an Excel engine) are the real levers. 4. Audit your graders. Grader blind spots create false negatives—if you tune skills against such a grader, you'll discard skills that are actually correct.
Broader Context
The paper extends the "evaluation blind-spot" principle from model evaluation to agent evaluation: average pass rate is the biggest blind spot, merging gains and regressions into one number. Capability isn't a pass rate—it's the paired gain/regression structure. The cross-stage conclusion also echoes "division of labor beats unification": use grounding-specific and verification-specific tools, and reserve process guidance for where it actually helps.
---
Paper: The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents Authors: Darshan Tank, Baran Nama (Sentient Labs) arXiv: 2607.22520 Code: Not yet released
FAQ
Q1: Who is this for? Practitioners, researchers, and students working on LLM agents, agent tooling, and evaluation.
Q2: What is the core takeaway? Skill libraries' net gains mask substantial regressions (553 gains vs. 324 regressions); evaluate paired outcomes, and target skills at grounding and verification rather than method.
Q3: Is there open-source code? Not yet—see the arXiv link above for the paper.