English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Neural Thickets: Random Guessing Can Match RL for Post-Training Large Language Models

Forum topic · 小凯 · 2026-03-19

Summary

A detailed Chinese-language walkthrough of a MIT paper titled 'Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights' (arXiv:2603.12228). The paper shows that in sufficiently pretrained large language models, random Gaussian perturbations of the weights frequently yield models that improve on downstream tasks—task experts are densely clustered around pretrained weights. Experiments on Qwen2.5 (0.5B–32B) show the fraction of perturbations matching or beating baseline on GSM8K rises from 0% to 64% with scale, and a 'Spectral Discordance' metric shows these solutions are diverse specialists rather than generalists. Based on this, the authors propose RandOpt, a fully parallel algorithm that samples N perturbations, selects the best K, and ensembles them via majority voting. On a 200-GPU GH200 cluster with Olmo-3-7B-Instruct, RandOpt reached 70% accuracy on Countdown in 3.2 minutes, competitive with PPO, GRPO, and ES. The post also covers limits: RandOpt requires pretrained weights, gains saturate, and K-model inference costs K-fold more (mitigated by distillation). A toy MLP experiment suggests pretrained models sit at a strategic point analogous to MAML-style good initializations or the Baldwin effect.

This post is a long-form Chinese-language explainer of the MIT paper *Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights* (arXiv:2603.12228). Below is a structured English summary of its contents.

Key points

  • From haystack to thicket: In 2001, Schmidhuber, Hochreiter, and Bengio argued that random guessing cannot be a reasonable learning algorithm—in a billion-dimensional parameter space, the probability of randomly hitting a capable solution is negligible. The MIT paper shows this holds only for small or untrained models.
  • Core finding: Around the weights of a large, well-pretrained language model, random Gaussian perturbations frequently match or improve downstream task performance. Task "experts" are dense around pretrained weights—a "neural thicket" rather than a needle in a haystack.
  • Scaling law for solution density: On Qwen2.5 models (0.5B–32B), the fraction of perturbations matching or exceeding baseline on GSM8K rose from 0% to 64% as scale increased; on Countdown it rose from 8% to 60%. Solution density grows with model size, pretraining data volume, and pretraining quality.
  • Specialists, not generalists: A proposed metric, Spectral Discordance, quantifies whether sampled solutions rank similarly across tasks. As scale grows, discordance increases—sampled solutions excel on specific tasks (math, chemistry, writing, code) and form distinct clusters in PCA projections. RGB landscape maps across a 2D parameter slice show task-wise performance regions that are largely independent.
  • The RandOpt algorithm

    1. Sample N Gaussian perturbations from the pretrained weights. 2. Evaluate them on training data (fully parallel). 3. Select the top K performers. 4. Ensemble at test time via majority voting.

  • Wall-clock speed: Unlike sequential methods (gradient descent, PPO, GRPO, ES), RandOpt has O(1) wall-clock training complexity. Using Olmo-3-7B-Instruct on Countdown with N=2000, K=50 on 200 GH200 GPUs, training took 3.2 minutes to reach 70% accuracy.
  • Competitiveness: Across Countdown, GSM8K, MATH-500, OlympiadBench, MBPP, ROCStories, and USPTO—on Qwen2.5, Llama, and OLMo3 families—RandOpt (K=50) performs comparably to PPO/GRPO/ES in most settings, sometimes better.
  • Ensembling matters: K=50 clearly beats K=1, since each perturbation is a narrow expert. Inference cost rises K-fold; distillation from the ensemble back into a single model can recover most performance at normal inference cost.
  • Why thickets emerge

    A minimal MLP experiment on 1D signal prediction (sine, linear, square, sawtooth waves) compared three regimes:

  • Needle-in-a-haystack (no pretraining): random perturbations are useless.
  • Thicket regime (mixed pretraining): perturbations yield diverse specialists, one per signal type, enabling good ensembled predictions.
  • Plateau regime (pretraining matched to the test distribution): the pretrained weights are already near-optimal and perturbation hurts.
  • The authors interpret pretraining not merely as knowledge acquisition but as positioning in parameter space—automatically producing a MAML-like good initialization, echoing the Baldwin effect in evolutionary biology.

    Limitations

  • Requires fully pretrained weights; useless from random initialization.
  • Gains appear to saturate with scale, N, and K—far-beyond-baseline improvements likely require leaving the thicket, where gradient-based search again dominates.
  • K-fold inference cost (mitigable via distillation).
  • The paper positions RandOpt as a complement to, not a replacement for, gradient descent.
  • Notable details

  • Decomposing GSM8K gains: RandOpt (K=50) reached 86.7% accuracy, with improvements split into "reasoning thicket" (12.3%—genuinely fixed reasoning) and "format thicket" (19.0%—fixed answer formatting).
  • Preliminary results (appendix J) suggest "color thickets" also exist in image generation models.
  • Implications discussed: rethinking the complexity of post-training methods, prioritizing pretraining quality, communication-free distributed training, federated learning applications, and AI-safety questions about unexpressed capabilities around pretrained weights.

References cited in the post

1. Gan, Y., et al. (2026). Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights. *arXiv:2603.12228*. Project page: https://thickets.mit.edu 2. Schmidhuber, J., Hochreiter, S., & Bengio, Y. (2001). Evaluating benchmark problems by random guessing. *NeurIPS*. 3. Finn, C., Abbeel, P., & Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. *ICML*. 4. Salimans, T., et al. (2017). Evolution strategies as a scalable alternative to reinforcement learning. *arXiv:1703.03864*. 5. Baldwin, J. M. (1896). A new factor in evolution. *The American Naturalist*, 30(354).

Tags

#large-language-models#post-training#randopt#random-search#mit#pretraining#ensembles#scaling-laws

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168913