This post is a long-form Chinese-language explainer of the MIT paper *Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights* (arXiv:2603.12228). Below is a structured English summary of its contents.
Key points
- From haystack to thicket: In 2001, Schmidhuber, Hochreiter, and Bengio argued that random guessing cannot be a reasonable learning algorithm—in a billion-dimensional parameter space, the probability of randomly hitting a capable solution is negligible. The MIT paper shows this holds only for small or untrained models.
- Core finding: Around the weights of a large, well-pretrained language model, random Gaussian perturbations frequently match or improve downstream task performance. Task "experts" are dense around pretrained weights—a "neural thicket" rather than a needle in a haystack.
- Scaling law for solution density: On Qwen2.5 models (0.5B–32B), the fraction of perturbations matching or exceeding baseline on GSM8K rose from 0% to 64% as scale increased; on Countdown it rose from 8% to 60%. Solution density grows with model size, pretraining data volume, and pretraining quality.
- Specialists, not generalists: A proposed metric, Spectral Discordance, quantifies whether sampled solutions rank similarly across tasks. As scale grows, discordance increases—sampled solutions excel on specific tasks (math, chemistry, writing, code) and form distinct clusters in PCA projections. RGB landscape maps across a 2D parameter slice show task-wise performance regions that are largely independent.
- Wall-clock speed: Unlike sequential methods (gradient descent, PPO, GRPO, ES), RandOpt has O(1) wall-clock training complexity. Using Olmo-3-7B-Instruct on Countdown with N=2000, K=50 on 200 GH200 GPUs, training took 3.2 minutes to reach 70% accuracy.
- Competitiveness: Across Countdown, GSM8K, MATH-500, OlympiadBench, MBPP, ROCStories, and USPTO—on Qwen2.5, Llama, and OLMo3 families—RandOpt (K=50) performs comparably to PPO/GRPO/ES in most settings, sometimes better.
- Ensembling matters: K=50 clearly beats K=1, since each perturbation is a narrow expert. Inference cost rises K-fold; distillation from the ensemble back into a single model can recover most performance at normal inference cost.
- Needle-in-a-haystack (no pretraining): random perturbations are useless.
- Thicket regime (mixed pretraining): perturbations yield diverse specialists, one per signal type, enabling good ensembled predictions.
- Plateau regime (pretraining matched to the test distribution): the pretrained weights are already near-optimal and perturbation hurts.
- Requires fully pretrained weights; useless from random initialization.
- Gains appear to saturate with scale, N, and K—far-beyond-baseline improvements likely require leaving the thicket, where gradient-based search again dominates.
- K-fold inference cost (mitigable via distillation).
- The paper positions RandOpt as a complement to, not a replacement for, gradient descent.
- Decomposing GSM8K gains: RandOpt (K=50) reached 86.7% accuracy, with improvements split into "reasoning thicket" (12.3%—genuinely fixed reasoning) and "format thicket" (19.0%—fixed answer formatting).
- Preliminary results (appendix J) suggest "color thickets" also exist in image generation models.
- Implications discussed: rethinking the complexity of post-training methods, prioritizing pretraining quality, communication-free distributed training, federated learning applications, and AI-safety questions about unexpressed capabilities around pretrained weights.
The RandOpt algorithm
1. Sample N Gaussian perturbations from the pretrained weights. 2. Evaluate them on training data (fully parallel). 3. Select the top K performers. 4. Ensemble at test time via majority voting.
Why thickets emerge
A minimal MLP experiment on 1D signal prediction (sine, linear, square, sawtooth waves) compared three regimes:
The authors interpret pretraining not merely as knowledge acquisition but as positioning in parameter space—automatically producing a MAML-like good initialization, echoing the Baldwin effect in evolutionary biology.
Limitations
Notable details
References cited in the post
1. Gan, Y., et al. (2026). Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights. *arXiv:2603.12228*. Project page: https://thickets.mit.edu 2. Schmidhuber, J., Hochreiter, S., & Bengio, Y. (2001). Evaluating benchmark problems by random guessing. *NeurIPS*. 3. Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. *ICML*. 4. Salimans, T., et al. (2017). Evolution strategies as a scalable alternative to reinforcement learning. *arXiv:1703.03864*. 5. Baldwin, J. M. (1896). A new factor in evolution. *The American Naturalist*, 30(354).