This post is a Feynman-style walkthrough of ST-EVO, a multi-agent framework in which the team's "seating chart" — the communication topology — evolves together with the speaking schedule as the conversation unfolds.
Three existing paradigms, and their limits
1. Static topology (AutoGen, LangChain fixed chains/star/graphs): the same fixed structure is used for every task, wasting communication on tasks that don't need it. 2. Spatial evolution (GPTSwarm, G-Designer): the topology is re-planned per task but stays frozen once the conversation starts. 3. Temporal evolution (AFlow, STEER): a coordinator decides when agents speak, but not who should communicate with whom — the topology is ignored.
ST-EVO's claim: topology and timing should co-evolve.
Core: a flow-matching scheduler
The heart of ST-EVO is a compact Scheduler built on flow matching. Unlike diffusion models or GANs that generate from noise, flow matching can smoothly deform one distribution into another from an arbitrary starting point — exactly what topology evolution needs: start from the *current* graph and adjust edges based on task state.
Architecture:
- GCN encodes the current communication topology
- MLP-based flow-matching network predicts the deformation "velocity" (how the graph should change next)
- Conditioning inputs: query embedding (task type) + iteration encoding (position in the conversation)
- Predictive Entropy (PE): how uncertain the model's predictions are
- VarEntropy (VE): variance of uncertainty across agents
- αᵢ: access frequency (higher is better)
- cᵢ: compute cost (expensive entries get evicted)
- uᵢ: uncertainty (unstable entries get evicted)
- Static systems (Chain, Star, Tree, Random Graph): performance drops 6.5%-12.7%
- Single-dimension evolution (G-Designer, STEER): drops <5%
- ST-EVO: 88.4% → 88.1%, essentially unchanged
After every round, the topology is recomputed: edges are added, removed, or rewired as the discussion progresses.
Four system states via entropy awareness
ST-EVO combines two signals:
| PE | VE | State | Response | |----|----|-------|----------| | Low | Low | High confidence | Prune the topology, cut redundant communication | | High | High | Conflict / disagreement | Add connections to let views collide | | High | Low | Collective ignorance | Introduce new information sources, deeper thinking | | Low | High | Overconfidence anomaly (rare, dangerous) | Possible agent "hijacking" the consensus — be cautious |
This feeds a stability reward used to train the scheduler: R_sta = -(α·PE + β·1/(1+VE)). Higher stability means a more certain, consistent team state.
Engineering detail: entropy is computed only over the top 10-20% highest-entropy tokens, since punctuation and stopwords are trivially high-confidence and dilute the signal.
Experience replay: learning from history
Each successful scheduling trajectory is stored (query embedding, latent topology sequence [L̂₁, ..., L̂_T], compute cost, uncertainty, access frequency). For a new task, the most similar past trajectories are retrieved and used as a regularization signal — guidance, not copying.
Eviction score: S(mᵢ) = (1 + log(αᵢ + 1)) / (cᵢ · |uᵢ| + ε)
Frequent, cheap, stable experiences survive; expensive and unreliable ones are cleaned up — smarter than plain LRU.
Benchmark results: where the 5-25% gains come from
| Benchmark | ST-EVO | Gain vs. baselines | |-----------|--------|--------------------| | MMLU | 89.85% | +9.38% | | GSM8K | 97.60% | +10.45% | | HumanEval | 94.36% | +21.08% | | DS-1000 | 58.65% | +20.25% | | DDXPlus | 82.50% | +26.10% | | AQuA | 86.56% | +17.29% |
The largest gain is on medical diagnosis (DDXPlus), which naturally requires multi-hop, multi-perspective collaboration (symptoms → differential diagnosis → tests → confirmation). On coding benchmarks, the tester agent can reconnect directly to the implementer agent upon finding a bug, instead of routing through fixed intermediate roles.
Token savings: nearly half
| Benchmark | ST-EVO | Best baseline | Savings | |-----------|--------|---------------|---------| | MMLU | 1.3M | 1.6M-1.9M | ~50% | | GSM8K | 1.8M | 1.9M-3.3M | ~55% | | HumanEval | 0.28M | 0.29M-0.38M | ~15-26% |
Savings come from sparse topologies: in many rounds a 3-4 agent graph has only 2-3 active edges instead of full connectivity. The entropy mechanism actively prunes the graph when the team is confident.
Counter-intuitive robustness under attack
Under prompt injection attacks:
One-sentence takeaway
> ST-EVO doesn't optimize *what* multi-agent systems say — it optimizes *the structure* of their conversation. Who talks to whom, when, and when to stay silent turns out to matter more than any individual utterance. Flow matching turns topology evolution from art into engineering; entropy awareness and experience replay let the team learn from its own chaos.
Paper: https://arxiv.org/abs/2602.14681