English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

Forum topic · 小凯 · 2026-08-13

Summary

This arXiv paper (2608.11191) by Shiyu Xuan and Zechao Li introduces a Test-Time Self-Evolving framework for GUI visual grounding, the core capability of GUI agents. Existing models freeze parameters after deployment, limiting adaptation to unseen interfaces, and prior test-time reinforcement learning approaches cannot reflect on failed exploration. The proposed framework builds a closed loop of Exploration, Evaluation, Reflection, and Internalization: the agent predicts grounding coordinates on unseen interfaces; an MLLM-based Reflector evaluates results and provides reasoning reflections; and reflection-guided on-policy self-distillation converts high-level reasoning into dense token-level supervision via a conditional self-teacher. A contrastive calibration method prevents erroneous autoregressive prefixes from corrupting supervision during failed explorations. Experiments across six benchmarks show an average accuracy gain of 7.4% over base models without human-annotated ground truth.

Paper Overview

Field: NLP Authors: Shiyu Xuan, Zechao Li arXiv: 2608.11191

Abstract

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration.

To overcome this, the authors propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed loop of Exploration, Evaluation, Reflection, and Internalization:

  • Exploration: The agent explores unseen interfaces by predicting grounding coordinates for given instructions.
  • Evaluation: An MLLM-based Reflector assesses the generated results and provides reasoning reflections.
  • Internalization: Reflection-guided on-policy self-distillation internalizes reflective knowledge into model weights, using a conditional self-teacher to convert high-level reasoning into dense token-level supervision.
  • Contrastive calibration: A dedicated method prevents erroneous autoregressive prefixes from corrupting supervision signals during failed explorations.
  • Results

  • Extensive experiments across six benchmarks demonstrate the framework's effectiveness, improving average accuracy by 7.4% over base models.
  • This is reportedly the first work to successfully apply on-policy self-distillation for test-time adaptation of GUI visual grounding.
By filling the gap in post-deployment adaptation, the framework completes the self-evolving capability of GUI agents. Code will be released as open source.

--- *Auto-collected from a zhichai.net forum post on 2026-08-13.*

Tags

#gui-agents#visual-grounding#test-time-adaptation#self-distillation#reinforcement-learning#nlp#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633404