> Paper: Never Stop Learning: A Survey of Continual Learning and Self-Iteration in Large Language Models > Author: Deli Chen (auto-generated by the Deli AutoResearch framework) > Models: DeepSeek-V4-Pro (text generation and reasoning) + GPT-Image-2 (figures) > Version: V5 (updated June 4, 2026)
This post is a Chinese tech-forum deep dive into an AI-generated survey paper. Note that claims attributed to the paper (benchmark numbers, citations dated 2026) come from an AI research framework and may include fabricated or predictive references — treat quantitative results with caution.
Key points
- Three-dimensional taxonomy: The paper claims to be the first framework jointly covering continual learning (CL) and self-improvement, organized along What (Knowledge / Skills / Alignment), How (External Signal / Self-Generated / Architectural), and When (Offline / Online / Test-time).
- Method mapping examples: EWC → Knowledge + External + Offline; SPIN → Skills + Self-Generated + Offline; o1 → Reasoning + Self-Generated + Test-time; RLHF → Alignment + External + Online; STaR → Skills + Self-Generated + Offline.
- Core insight: CL and self-improvement, previously separate research communities, can be understood under one unified framework.
- Forgetting decreases with scale: models >10B forget 30–50% less than 1B models.
- Regularization methods become relatively ineffective at high dimensionality.
- Replay ratio needed stays roughly constant from 1B to 70B.
- LoRA adapters become proportionally cheaper as models grow.
- AI-generated paper limitations: some citations may be predictive/fabricated; computational cost discussion (e.g., inference routing overhead of LoRA isolation) is shallow; multimodal continual learning coverage is thin.
- Taxonomy boundaries blur in practice — hybrid methods like LoRA + replay span multiple quadrants.
- Self-improvement optimism: verifier quality is rarely guaranteed in open domains where "better output" lacks consensus.
- Scale effects causality: the paper doesn't distinguish whether less forgetting at scale is a capacity property or a training-dynamics effect.
- Short term (1–2 yrs): optimal LoRA/replay mixing ratios, dynamic verifier design, efficient test-time self-improvement.
- Mid term (3–5 yrs): recursive self-improvement (improving the learning method itself), lifelong knowledge graphs, real-time alignment drift monitoring.
- Long term (5+ yrs): theoretical limits of bootstrapping; autonomous scientific discovery.
Five method families reviewed (100+ papers)
| Family | Mechanism | Representatives | Forgetting prevention | LLM practicality | |---|---|---|---|---| | Parameter isolation | Per-task parameter subsets | Progressive Networks, LoRA, AdapterFusion | Strongest | Highest — LoRA scales well on 100B+ models | | Regularization | Auxiliary losses penalize parameter drift | EWC, SI, LwF | Weak | Poor above 10B (Fisher diagonal approximation degrades) | | Replay | Mix stored/generated old samples | GEM, DER++, generative replay | Moderate | Moderate — 1–5% replay ratio stable across scales | | Architectural | MoE, modular designs | Switch Transformer, CP-MoE | Moderate | Moderate — inference routing overhead | | Prompt-based | Learnable prompts, frozen backbone | L2P, DualPrompt, CODA-Prompt | Theoretically zero | Moderate — good for fast task switching |
Reported benchmark results (TRACE, Llama-2 7B)
| Method | Avg. accuracy (AA) | Backward transfer (BWT) | |---|---|---| | Sequential fine-tuning (no CL) | 52.3% | -31.4 | | EWC | 64.8% | -18.2 | | Replay (5%) | 71.2% | -8.7 | | LoRA isolation | 73.5% | -3.1 | | Joint training (upper bound) | 83.7% | 0.0 |
LoRA isolation reportedly reduces catastrophic forgetting from -31.4 to -3.1, nearly eliminating it.
Scale effects
Recommended practice for 100B+ models: LoRA isolation for new domain adaptation + lightweight replay (1–2% original data) for general knowledge retention.
Self-improvement theory
The paper formalizes an iterative refinement loop:
Convergence requirement: the quality of generated training signals must exceed the average output quality of the current policy — otherwise iteration collapses. Perfect verifiers (e.g., game outcomes) guarantee convergence; ambiguous verifiers (e.g., text quality) need additional mechanisms such as consistency filtering, complexity selection, or execution-based verification.
Relation to the companion piece
The post connects this survey to a previous one, "From Copilots to Colleagues": the first answers "what can AI do?" (L1–L5 autonomy levels), this one answers "how does AI maintain and grow capability?" (What-How-When taxonomy). The recursive twist: an L4-level self-improving system produced a survey about self-improvement.
Six open challenges
1. Catastrophic forgetting at 100B+ scale (scale-aware regularization theory) 2. Reward hacking (formal verification, constitutional constraints) 3. Evaluation under distribution shift (dynamic benchmarks) 4. Stability–plasticity of alignment (multi-objective Pareto optimization) 5. Safe continual alignment (real-time monitoring of value drift) 6. Theoretical limits of self-improvement (complexity-theoretic analysis of pure bootstrapping)