> One-line verdict: The paper's solid contribution is a clean experimental demonstration that naive uniform LoRA falls far short of full fine-tuning on multi-step procedures; its overreach lies in elevating this empirical negative result into a theorem-like claim that procedural knowledge is fundamentally non-low-rank. We should trust the behavioral findings and be cautious about the title claim.
---
1. What is this paper?
- Title: *Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures*
- Authors: Simon Dennis, Kevin Shabahang, Hao Guo, Rivaan Patil (University of Melbourne)
- ID: arXiv:2607.21612 (v1, 14 pages, submitted July 2026)
- Core question: Parameter-efficient fine-tuning (PEFT) methods like LoRA have become the default for fine-tuning, succeeding repeatedly on instruction following, style transfer, and factual adaptation. But does LoRA still match full fine-tuning for procedural knowledge — the ability to navigate conditional branches across multiple steps to a terminal state?
- Full fine-tuning is like rewriting the entire cookbook — every step and heat level rewritten into muscle memory.
- LoRA merely adds two margin notes (ΔW = B·A with tiny rank r), hoping those notes can retune the whole book.
- Instruction following and style transfer are 'local, smooth' flavors — margin notes suffice.
- Procedural knowledge is a dance of 'choose a branch first, decide the next step from prior state, loop until termination' — it spans the whole book and is interdependent; two margin notes cannot capture it.
- Completion rate measures surface token distribution; task success measures the state→action mapping — the two are decoupled. Low-rank B·A suffices to memorize frequent n-grams and template utterances, but struggles to encode conditional transitions in the implicit state space (effective rank 888+).
- Root of high-rank degradation: decision trees are sparse, local, and conditional (86–2381 paths, nested loops); low-rank spaces express smooth, globally correlated perturbations. As r grows, capacity is spent fitting specific phrasings of particular paths — overfitting surfaces, not abstract transition rules.
- MLP is the true bottleneck: the state→action mapping lives in the MLP, which has the highest effective rank (down_proj 1872) and the largest capacity gap; K/V appearing 'adaptive' is purely because GQA compresses their maximum rank to hundreds — a dimensional artifact, not evidence that K/V carries procedural knowledge.
- Only uniform-rank naive LoRA was tested; heterogeneous-rank (AdaLoRA-style) methods were not.
- The title's 'not low-rank / fundamental' overstates the evidence. The load-bearing assumption of the SVD argument is that full-FT ΔW is the 'ground truth' LoRA must approximate; but Eckart–Young only bounds approximation of that particular ΔW — it does not prove no low-rank update exists that reproduces the behavior. The paper's Limitations section concedes this.
- Free/hidden parameters: α=2r, LR 2e-4 vs 2e-5 (LoRA gets 10× and was not tuned per-rank), effective rank threshold never defined.
- Judges are LLMs, no human validation; only the Qwen family, only three customer-service domains.
- Untested variants: AdaLoRA/heterogeneous ranks, LoRA+, DoRA, VeRA, QLoRA, per-rank tuning — if any of these approaches full FT at r≤128, the 'structurally infeasible' claim is refuted.
- The 'surface fluency' trap: mainstream agents (LoRA-based models + external prompt/tool orchestration) delegate procedural correctness to prompts and scheduling; whenever the model must internalize conditional branches and cross-turn state, LoRA's rank limit of 128 captures only 43–51% of the energy and will take wrong turns. Failure modes map one-to-one: hallucinated steps ↔ overfit conversational patterns; lost state ↔ state loss resurging in later turns; infinite loops ↔ MLP effective rank 1872 covered only 7–13%.
- Pragmatic trade-offs: for multi-step workflows with conditional branches and cross-turn state (≥14 nodes, dense branching), the core model warrants full FT (3B runs on a single card; 8B needs 1–2× A100; one-time training, weight-level 'compilation', inference cost two orders of magnitude lower than orchestrated agents). For pure instruction/style/single-turn retrieval, or compute-constrained settings, keep LoRA + external state machine.
- Evaluation metrics: abandon the misleading proxy of 'dialogue completion rate'; track terminal-state reach rate, branch coverage, state-tracking accuracy, step-level correctness instead.
- Five practical rules: ① decouple routing from generation (a small fully-FT model handles state and routing, LoRA handles style); ② externalize the state machine (write workflows as explicit graphs/DSL; the model outputs intents, not direct actions); ③ specialize small models (full FT for procedural types, LoRA for conversational ones); ④ regression-test conditional branches (pre-deploy LLM-as-user multi-turn simulation; block on failure); ⑤ lock versions and record training topology.
- arXiv:2607.21612 | University of Melbourne | 14 pages | submitted July 2026
- Task Success: LoRA ≤2.54 vs Full FT 4.11 (p<0.001)
- Dialogue completion rate: LoRA 95.5–99.0% vs 100%
- Effective rank: Travel 888 / Zoom 761 / Insurance 1026 (MLP down_proj up to 1872)
- Energy captured at r=128: 42.5% / 50.8% / 43.0% (only 43–51%)
- Cross-domain gap: 0.8–2.2 points, largest on Insurance
---
2. Intuition: what is 'low-rank', what is 'procedural knowledge'
Mathematically: full fine-tuning weight updates ΔW have an average effective rank of 761–1026 (the Insurance MLP down_proj reaches 1872); LoRA even at r=128 can, by the Eckart–Young theorem, capture at most 43–51% of the squared Frobenius energy. A large gap.
---
3. What the experiments show (key numbers)
Three task domains: Travel booking (14 nodes), Zoom support (14 nodes), Insurance claims (55 nodes, nearly 4× complexity). Models: Qwen2.5-3B / Qwen3-8B. Scoring: Claude Sonnet 4.5 as primary judge + GPT-4.1 as secondary (dual judges).
| Metric | LoRA (r=16–128) | Full FT | Significance | |---|---|---|---| | Task Success (out of 5) | ≤ 2.54 | 4.11 | p<0.001, \|d\|>1.9 | | Dialogue completion rate | 95.5–99.0% | 100% | — | | Cross-domain gap (mean of r=32/128) | behind by 0.8–2.2 points | — | Largest gap on Insurance |
The paradox is stark: LoRA speaks fluently (completion rate near-perfect) but walks the wrong path (success rate barely half). The authors call this *completion without correctness*. Stranger still: higher rank (>32) scores drop (2.54→2.44→2.10); §3.4 shows LoRA's held-out loss is actually lower (0.81 vs 0.89) while behavior is worse — 'LoRA overfits without learning': it memorizes what the dialogue looks like without learning why the steps are taken.
---
4. Three-way analysis (integrating three perspectives)
Perspective A · Mechanism: why the low-rank bottleneck breaks on procedural knowledge
Perspective B · Reviewer: the title is too bold; mechanistic extrapolation overextended
The behavioral results are credible (systematic ablations + cross-domain replication + dual judges; rated by Pith as a *clean, practical negative result*), but the scope is narrow:
Perspective C · Architecture: beyond surface fluency — what to do
---
5. Final conclusions
1. Trust the behavior, doubt the title. The paper's greatest value is a clean, reproducible empirical proof that naive uniform LoRA systematically underperforms full FT on multi-step procedures, with failure precisely localized at the behavior layer rather than the loss layer. This finding should be adopted. 2. Do not treat it as iron law. The claim of 'procedural knowledge is non-low-rank / a fundamental limitation' is under-evidenced; heterogeneous-rank variants, DoRA, and others may yet reverse it. Cite it as an 'empirical negative result', not an 'impossibility proof'. 3. A lesson for agent architectures: whenever core workflows must 'walk the right path', do not stake correctness on naive LoRA weights — either use full FT or externalize routing/state. Deceiving yourself with 'fluent dialogue' is the biggest trap this paper exposes. 4. Reform evaluation: replace 'didn't crash = success' completion rates with terminal-state reach rate, branch coverage, and cross-turn state accuracy.
---