Paper Overview
Field: NLP Authors: Shiyu Xuan, Zechao Li arXiv: 2608.11191
Abstract
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration.
To overcome this, the authors propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed loop of Exploration, Evaluation, Reflection, and Internalization:
- Exploration: The agent explores unseen interfaces by predicting grounding coordinates for given instructions.
- Evaluation: An MLLM-based Reflector assesses the generated results and provides reasoning reflections.
- Internalization: Reflection-guided on-policy self-distillation internalizes reflective knowledge into model weights, using a conditional self-teacher to convert high-level reasoning into dense token-level supervision.
- Contrastive calibration: A dedicated method prevents erroneous autoregressive prefixes from corrupting supervision signals during failed explorations.
- Extensive experiments across six benchmarks demonstrate the framework's effectiveness, improving average accuracy by 7.4% over base models.
- This is reportedly the first work to successfully apply on-policy self-distillation for test-time adaptation of GUI visual grounding.
Results
--- *Auto-collected from a zhichai.net forum post on 2026-08-13.*