English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ST-EVO: Jointly Evolving Communication Topology and Timing in Multi-Agent LLM Systems

Forum topic · 小凯 · 2026-06-29

Summary

ST-EVO is a multi-agent LLM framework that evolves both the communication topology (who talks to whom) and the temporal scheduling (who speaks when) during collaborative problem solving. Its core component is a compact flow-matching scheduler: a GCN encodes the current communication graph, and an MLP-based flow-matching network predicts the next, better-suited topology conditioned on the query embedding and iteration index. An entropy-aware mechanism combines predictive entropy (PE) and variance across agents (VE) to classify system states (high confidence, conflict, collective ignorance, overconfidence) and guide a stability reward during training. A retrieval-augmented experience memory stores successful topology trajectories with a retention score favoring frequent, cheap, stable entries. Across nine benchmarks, ST-EVO reports gains of roughly 5-25%, including +26.10% on DDXPlus medical diagnosis, 94.36% on HumanEval, and 97.60% on GSM8K, while cutting token consumption by ~15-55% versus the best baselines. Under prompt-injection attacks, accuracy barely changes (88.4% to 88.1%), whereas static topologies drop 6.5-12.7%, suggesting dynamic rewiring provides fault isolation. Paper: https://arxiv.org/abs/2602.14681

This post is a Feynman-style walkthrough of ST-EVO, a multi-agent framework in which the team's "seating chart" — the communication topology — evolves together with the speaking schedule as the conversation unfolds.

Three existing paradigms, and their limits

1. Static topology (AutoGen, LangChain fixed chains/star/graphs): the same fixed structure is used for every task, wasting communication on tasks that don't need it. 2. Spatial evolution (GPTSwarm, G-Designer): the topology is re-planned per task but stays frozen once the conversation starts. 3. Temporal evolution (AFlow, STEER): a coordinator decides when agents speak, but not who should communicate with whom — the topology is ignored.

ST-EVO's claim: topology and timing should co-evolve.

Core: a flow-matching scheduler

The heart of ST-EVO is a compact Scheduler built on flow matching. Unlike diffusion models or GANs that generate from noise, flow matching can smoothly deform one distribution into another from an arbitrary starting point — exactly what topology evolution needs: start from the *current* graph and adjust edges based on task state.

Architecture:

  • GCN encodes the current communication topology
  • MLP-based flow-matching network predicts the deformation "velocity" (how the graph should change next)
  • Conditioning inputs: query embedding (task type) + iteration encoding (position in the conversation)
  • After every round, the topology is recomputed: edges are added, removed, or rewired as the discussion progresses.

    Four system states via entropy awareness

    ST-EVO combines two signals:

  • Predictive Entropy (PE): how uncertain the model's predictions are
  • VarEntropy (VE): variance of uncertainty across agents
  • | PE | VE | State | Response | |----|----|-------|----------| | Low | Low | High confidence | Prune the topology, cut redundant communication | | High | High | Conflict / disagreement | Add connections to let views collide | | High | Low | Collective ignorance | Introduce new information sources, deeper thinking | | Low | High | Overconfidence anomaly (rare, dangerous) | Possible agent "hijacking" the consensus — be cautious |

    This feeds a stability reward used to train the scheduler: R_sta = -(α·PE + β·1/(1+VE)). Higher stability means a more certain, consistent team state.

    Engineering detail: entropy is computed only over the top 10-20% highest-entropy tokens, since punctuation and stopwords are trivially high-confidence and dilute the signal.

    Experience replay: learning from history

    Each successful scheduling trajectory is stored (query embedding, latent topology sequence [L̂₁, ..., L̂_T], compute cost, uncertainty, access frequency). For a new task, the most similar past trajectories are retrieved and used as a regularization signal — guidance, not copying.

    Eviction score: S(mᵢ) = (1 + log(αᵢ + 1)) / (cᵢ · |uᵢ| + ε)

  • αᵢ: access frequency (higher is better)
  • cᵢ: compute cost (expensive entries get evicted)
  • uᵢ: uncertainty (unstable entries get evicted)
  • Frequent, cheap, stable experiences survive; expensive and unreliable ones are cleaned up — smarter than plain LRU.

    Benchmark results: where the 5-25% gains come from

    | Benchmark | ST-EVO | Gain vs. baselines | |-----------|--------|--------------------| | MMLU | 89.85% | +9.38% | | GSM8K | 97.60% | +10.45% | | HumanEval | 94.36% | +21.08% | | DS-1000 | 58.65% | +20.25% | | DDXPlus | 82.50% | +26.10% | | AQuA | 86.56% | +17.29% |

    The largest gain is on medical diagnosis (DDXPlus), which naturally requires multi-hop, multi-perspective collaboration (symptoms → differential diagnosis → tests → confirmation). On coding benchmarks, the tester agent can reconnect directly to the implementer agent upon finding a bug, instead of routing through fixed intermediate roles.

    Token savings: nearly half

    | Benchmark | ST-EVO | Best baseline | Savings | |-----------|--------|---------------|---------| | MMLU | 1.3M | 1.6M-1.9M | ~50% | | GSM8K | 1.8M | 1.9M-3.3M | ~55% | | HumanEval | 0.28M | 0.29M-0.38M | ~15-26% |

    Savings come from sparse topologies: in many rounds a 3-4 agent graph has only 2-3 active edges instead of full connectivity. The entropy mechanism actively prunes the graph when the team is confident.

    Counter-intuitive robustness under attack

    Under prompt injection attacks:

  • Static systems (Chain, Star, Tree, Random Graph): performance drops 6.5%-12.7%
  • Single-dimension evolution (G-Designer, STEER): drops <5%
  • ST-EVO: 88.4% → 88.1%, essentially unchanged
Why? Attacks typically target specific agents. When one is poisoned, the entropy mechanism detects the conflict state (high PE, high VE) and the scheduler rewires the topology to bypass the compromised node — a natural form of fault isolation that static topologies lack.

One-sentence takeaway

> ST-EVO doesn't optimize *what* multi-agent systems say — it optimizes *the structure* of their conversation. Who talks to whom, when, and when to stay silent turns out to matter more than any individual utterance. Flow matching turns topology evolution from art into engineering; entropy awareness and experience replay let the team learn from its own chaos.

Paper: https://arxiv.org/abs/2602.14681

Tags

#st-evo#multi-agent-systems#llm#flow-matching#topology-evolution#entropy#prompt-injection#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208272