English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

EEVEE: Test-time Prompt Learning for LLM Agents in Real-world Task Streams

Forum topic · 小凯 · 2026-06-11

Summary

EEVEE is the first multi-dataset test-time prompt learning framework for LLM agents, enabling prompt optimization under real-world task streams. Existing methods are largely designed for single-dataset settings, whereas real applications require handling heterogeneous inputs from multiple datasets, domains, and task distributions. To mitigate cross-dataset interference, EEVEE introduces a router that partitions incoming inputs into task clusters and assigns them to suitable prompt configurations. The design is optimized via a router-prompt co-evolution strategy using interleaved router and prompt learning phases to address their mutual dependency. Experiments show average multi-benchmark improvements of 10.38 and 24.32 points over Qwen3-4B-Instruct and DeepSeek-V3.2 respectively, outperforming state-of-the-art methods GEPA and ACE by 37.2% and 48.2%. Paper: arXiv:2606.11182 by Weixian Xu, Shilong Liu, and Mengdi Wang.

Paper Overview

Research area: Machine Learning Authors: Weixian Xu, Shilong Liu, Mengdi Wang Published: 2026-06-09 arXiv: 2606.11182

Abstract

In this paper, the authors propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams. Existing methods are largely designed for single-dataset settings, while real-world applications require models to handle heterogeneous input streams drawn from multiple datasets, domains, and task distributions, limiting their practical applicability.

To mitigate cross-dataset interference, EEVEE introduces a router that partitions incoming inputs into task clusters and assigns them to suitable prompt configurations. This design is optimized via a router-prompt co-evolution strategy, which employs interleaved router and prompt learning phases to address their mutual dependency.

Key Results

  • Average multi-benchmark score improvements of 10.38 and 24.32 points over Qwen3-4B-Instruct and DeepSeek-V3.2, respectively
  • Outperforms state-of-the-art methods GEPA by 37.2% and ACE by 48.2%
  • Links

  • Paper: https://arxiv.org/abs/2606.11182
*Auto-collected on 2026-06-11.*

Tags

#llm-agents#test-time-prompt-learning#prompt-engineering#machine-learning#arxiv#router#multi-dataset

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981076