A Lost Correct Answer
EvoMap's internal experiment: a main agent decomposes tasks, and 30 sub-agents solve and report on 563 problems. The sub-agents actually got 373 correct. Yet the final delivered answer contained only 217 correct ones.
166 once-correct answers vanished in transmission — a 55.5% retention rate for intermediate correct answers.
This is not a capability problem. The sub-agents could solve the problems; the main agent could read the reports. But the main agent's context is finite: with 30 reports stuffed in, it must compress and discard what it deems unimportant. The loss is not "can't solve" but information loss from organizational structure.
Three Organizational Schemes, Three Accuracy Levels
Same model, same 563 tasks (100 logic, 250 general math, 63 competition math, 150 physics), three setups:
| Scheme | Description | Accuracy | |--------|-------------|----------| | Monolithic (Agent Loop) | All tasks in one context, one agent | 26.3% | | Master-Slave (Agent Team) | Main agent decomposes; sub-agents report; main agent merges 30 reports | 38.5% | | Swarm (Agent Swarm) | Atomic tasks, one agent per problem, answers written to agreed locations, program collects by ID | 70.7% |
A caveat: stuffing 563 unrelated problems into one context causes interference by design; one-agent-per-problem naturally avoids this. So the comparison mostly shows "parallelize when appropriate, and use a program to merge instead of having a model paraphrase" — not that swarm architecture is inherently superior.
The interesting data is the middle layer: 373 correct sub-agent answers reduced to 217 after the main agent's synthesis.
Channel Capacity: Where 166 Answers Went
The main agent's context is a finite-bandwidth semantic channel. Reports passing through it are compressed, traded off, and dropped — lossy transmission in the information-theoretic sense. The swarm replaces the semantic channel with a symbolic channel: answers go to agreed locations and a program collects them by task ID — lossless. This is the same principle as Euclid-MCP's "let the LLM be the poet, let Prolog be the accountant":
> Division of labor beats unification. Parallelize what should be parallel; merge programmatically rather than letting the model paraphrase.
The master-slave failure isn't archaic architecture — it's making the model do two things it's bad at simultaneously: solving subtasks and lossily summarizing others' subtasks.
Continual Learning: Update Parameters, or Update Memory?
EvoMap's real question is continual learning for models with frozen parameters. Candidate routes:
- LoRA: low-rank adapters — changes the model's "instincts" (like editing DNA).
- AI4AI: AI-driven research loops.
- EvoMap: experience crystallized into Genes (structured experience records in
genes.json) — changes "experience" (like RNA editing). - With only social information (who collaborated with whom), agents pick "friends of friends" — cliques.
- With task information (expertise, accuracy), selection shifts to "who can get it done" — high scorers become hubs.
- EvoMap/EvoX: gitee.com/EvoMap
- AgentNet: arXiv:2504.00587
- OpenAI Swarm (archived): github.com/openai/swarm
- Kimi K3 Agent Swarm: moonshot.ai
These aren't mutually exclusive; they operate at matched granularities:
| Level | Mechanism | Analogy | Granularity | |-------|-----------|---------|-------------| | Parameter | LoRA | DNA editing | Parameter-level | | Memory | Gene/Memory | RNA editing | Behavior-level | | Architecture | Agent organization | Population structure | System-level |
EvoMap chose the memory layer because, under the "no retraining" constraint, it offers the best cost-benefit — hence their path from Evolver (plugin) to EvoMap (platform) to EvoX (swarm) stays within memory + architecture layers.
Gene vs SKILL.md: Executable vs Readable Knowledge
EvoMap's GEP protocol is a six-step loop: Scan (watch logs for errors/stalls), Signal (normalize to signals), Intent (plan evolution direction), Mutate (generate new code/prompt strategies), Validate (sandbox tests), Solidify (write verified capabilities into genes.json).
Key difference: SKILL.md is documentation written for humans that models happen to read; Gene keeps only what changes agent behavior.
The author's reservation: an optimization checklist without a reflection log tends toward local optima. Human expertise includes "what not to do" and "why" — if Genes keep only what works, agents can't trace back why previous approaches failed. This mirrors the finding of *Looping Is Not Reliability*: previously correct does not mean currently correct. When the environment shifts, an agent with only success records has no reflection log to consult.
The Real Fork: Oligopoly Risk
EvoMap's second experiment: 24 identical agents solve problems and deposit a Gene per task; experience accumulates into specialization. Then some ties in their network are randomly cut, and agents re-choose teammates:
The organization that emerges mirrors the information the system exposes. Academic work points the same way: AgentNet (arXiv:2504.00587) shows agents spontaneously differentiating in evolving networks.
But an unstated structural fork: will an outcome-oriented collaboration network self-reinforce into an oligopoly? High-scoring hubs get more tasks, more experience, higher scores; low-scoring agents never accumulate experience — a textbook Matthew effect. Real bee colonies don't allocate by historical accuracy but by proximity and discovery.
EvoMap's GEP protocol itself is centrally designed — even "choosing teammates" is protocol-specified. Strictly speaking, this is not "collective intelligence without a central brain" but "collective intelligence with protocol-before-emergence." Not a criticism: biological swarm protocols (the waggle dance, pheromone trails) are also gene-encoded, not emergent. The open question:
> If outcome-oriented networks self-reinforce into oligopolies, does EvoMap need an anti-Matthew mechanism? Or is oligopoly the correct end state of continual learning — a few ever-more-specialized expert agents with the rest degenerating into routers?
Possible answers: oligopoly is fine (mirroring worker/scout bee differentiation), or it's a trap requiring perturbation mechanisms (random disturbance, forced rotation, experience decay) to preserve diversity.
Protocol Before Emergence
GEP's six steps are fixed. If agents someday need a seventh step — "Reflect" — who decides? If the protocol itself cannot evolve, continual learning only happens within the protocol's frame. A truly self-evolving agent cluster might need the protocol itself to become a Gene. But then: who writes the protocol for the protocol? Doesn't that require a central brain again?
Conclusion
Swarm intelligence doesn't eliminate the central brain — it moves it from runtime to design time. GEP is the central brain, defining the possibility space of runtime. This parallels Gubernaut's approach of moving regulation out of model weights into architecture: both solve problems by changing layers.
The true end state of continual learning may be three layers — design-time protocol, runtime experience, architectural organization — each learning continuously at its own granularity. LoRA changes parameters, Genes change experience, swarms change organization. What does the protocol change into?
References and projects: