English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldEvolver: Self-Evolving World Models for LLM Agent Planning

Forum topic · 小凯 · 2026-07-01

Summary

WorldEvolver is a self-evolving world model framework for long-horizon LLM agents introduced by Xuan Zhang, Wenxuan Zhang, and See-Kiong Ng (arXiv:2507.00003, July 2026). The framework improves agent foresight—predictions of action consequences before execution—by revising its deployment-time context while keeping the downstream agent and all model parameters frozen. WorldEvolver integrates three modules: Episodic Memory, which exploits real action transitions via retrieval-based simulation; Semantic Memory, which distills persistent heuristic rules from prediction-observation mismatches; and Selective Foresight, which filters low-confidence predictions before injecting them into the agent's reasoning context. Evaluated on ALFWorld and ScienceWorld, with world model prediction accuracy measured on Word2World and downstream success rates on AgentBoard, WorldEvolver achieves the highest prediction accuracy across three backbone models and outperforms other world model baselines in downstream agent success, showing that test-time memory revision enhances both prediction fidelity and planning performance.

Paper Overview

Research Area: Agent Authors: Xuan Zhang, Wenxuan Zhang, See-Kiong Ng Published: 2026-07-01 arXiv: 2507.00003

Abstract (translated)

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-making. In this paper, the authors introduce WorldEvolver, a self-evolving world model framework that revises its deployment-time context while keeping the downstream agent and all model parameters frozen.

WorldEvolver integrates three modules:

1. Episodic Memory — exploits real action transitions through retrieval-based simulation. 2. Semantic Memory — extracts persistent heuristic rules from prediction-observation mismatches. 3. Selective Foresight — filters low-confidence predictions before integrating them into the agent's reasoning context.

Evaluation

  • World model prediction accuracy: measured on Word2World
  • Downstream agent success rate: measured on AgentBoard
  • Evaluated on ALFWorld and ScienceWorld
Extensive experiments show that WorldEvolver achieves the highest prediction accuracy across three backbone models and leads other world model baselines in downstream agent success rate, demonstrating that test-time memory revision enhances both prediction fidelity and planning performance.

---

*Auto-collected on 2026-07-01. Source: zhichai.net forum post.*

Tags

#llm-agents#world-models#planning#memory#retrieval#arxiv-paper#agent-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208338