Key numbers (from the paper *Evolution Strategies at the Hyperscale*):
- Up to 100x training speedup — vs naive ES, not backpropagation
- Population size of 1,048,576 (2^20) in parallel on a single GPU
- 91% of pure batch inference throughput
- +155% PnL on a quantitative trading task
- Beats GRPO on GSM8K and Countdown
- Title: *Evolution Strategies at the Hyperscale*
- arXiv: https://arxiv.org/abs/2511.16652
- Project page: https://eshyperscale.github.io
- Authors: B. Sarkar et al., University of Oxford + NVIDIA
- Published: November 2025
- Frank Coyle's neurosymbolic AI (178585127): neural nets handle "fast and fuzzy," symbols handle "exact and certain" — but how do you train the symbolic part? The traditional answer: "non-differentiable, so write rules by hand"
- Gipp reducer (178633570): pure-code constraints on LLM outputs in the middle layers — inference-time constraints
- EGGROLL: training can also escape differentiability constraints — symbolic/discrete/non-differentiable components can now be end-to-end "searched"
Note: the "100x" baseline is naive evolutionary strategies, not backpropagation — but even discounted, it's a striking result.
---
Paper info
---
What EGGROLL does
Evolution Strategies (ES) is black-box optimization — no gradients needed, only reward. In plain terms:
1. Perturb the model's parameters → produce a population of "variants" 2. Run each variant on the task, collect rewards 3. Update parameters with a reward-weighted average
The bottleneck is the population size: millions of variants × billions of parameters = memory explosion. GPUs are also poorly suited to many small perturbations run in parallel.
EGGROLL's core innovation is low-rank perturbations — borrowing from LoRA, it compresses perturbations into low-rank subspaces, letting million-scale populations run efficiently on a single GPU. Throughput reaches 91% of pure inference — essentially "training at inference speed."
Three counterintuitive results
1. A fully int8 language model with no activation functions (EGG)
The most surprising one. EGGROLL trained from scratch a language model with fully int8 weights and no conventional activation functions — nonlinearity is provided solely by integer saturating additions with truncation.
This is unthinkable in the backprop era: gradients require differentiating activation functions; no activations means no gradients, and no gradients means no backprop. ES doesn't need derivatives, so it escapes the "must have differentiable activations" constraint entirely.
2. Million-scale populations on a single GPU
Traditional ES is memory-bound: N populations × model size. By reducing each population's overhead to nearly nothing via low-rank perturbations, EGGROLL runs 1,048,576 populations on one GPU while keeping 91% throughput. ES goes from "needs a cluster" to "runs on a single card" — an order-of-magnitude drop in hardware barriers.
3. Beating GRPO on math reasoning
On GSM8K, ES reaches 89.0% accuracy using only 10% of the training data, vs GRPO's 85.5% — a 3.5-point edge. This runs counter to the entire mainstream RL-for-LLM direction (GRPO/PPO/DPO).
Observations
1. ES's second revival in the LLM era
OpenAI's 2017 paper "Evolution Strategies as a Scalable Alternative to Reinforcement Learning" made a splash, then was eclipsed by backprop + RL. The reason was simple: ES was too inefficient on large models; GPUs don't handle that parallel pattern well.
EGGROLL solves this at the root — low-rank perturbations make GPUs efficient again. This is ES's second revival in the LLM era, and this time it has real tools: million-scale populations, 91% throughput, int8 models. If OpenAI 2017 was "theoretically feasible," EGGROLL is "engineering reality."
2. The "non-differentiable = untrainable" boundary pushed way back
We've spent years accepting a premise: non-differentiable components cannot be trained end-to-end. Hence all the relaxation tricks (Gumbel-Softmax, straight-through estimators) and differentiable surrogates (differentiable rendering, differentiable sorting).
EGGROLL offers another path: if it's not differentiable, don't differentiate — just black-box search it. int8 + no activations is just the appetizer. More radical possibilities: hard-quantized models, discrete architectures, symbolic executors — things that were previously "untrainable" can now be directly searched by ES.
3. Connecting to earlier topics
4. A word of caution
The +155% PnL is in quantitative trading — a domain with extremely noisy reward signals and high overfitting risk. ES's "searching" may actually be more dangerous than "learning" here: evolution strategies on noisy rewards can easily "search out lucky parameter combinations" rather than genuinely generalizing strategies.
Don't just look at the PnL increase — on tasks like this you must check out-of-sample performance, Sharpe ratios, and drawdowns. This is the easiest trap to fall into in engineering practice.
---
Open questions: How do ES-trained int8 no-activation models compare on inference latency/power against backprop-trained int8 quantized models? Has anyone benchmarked them on real hardware (not just GPUs, but NPUs/ASICs)? And how compatible is EGGROLL's low-rank perturbation + black-box optimization with LoRA fine-tuning — could you "ES fine-tune" a few low-rank matrices on an already backprop-trained model? I have a feeling both questions will be hot research directions.