English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics Curves

Forum topic · 小凯 · 2026-06-10

Summary

OmniGameArena (arXiv 2506.04867) is a real-time benchmark of twelve newly built Unreal Engine 5 games designed to evaluate vision-language model (VLM) agents. It covers Solo (7 games), PvP (3), and Coop (2) modes with unified action interfaces, enabling fair comparison of heterogeneous agent classes—commercial VLMs, open-weight VLMs, and specialized game policies—under the same protocol. The authors, Mingxian Lin, Shengju Qian, and Yuqi Liu, also introduce the Improvement Dynamics Curve (IDC), an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds. Beyond cold-start leaderboard scores, IDC exposes two additional observable metrics per (agent, game) pair: how scores evolve across reflection rounds, and how learned skills transfer to held-out task variants. The paper reports results for 12 VLM agents on the cold-start leaderboard and 4 top agents under IDC.

Paper Overview

  • Research Area: Computer Vision (CV)
  • Authors: Mingxian Lin, Shengju Qian, Yuqi Liu
  • Published: 2025-06-06
  • arXiv: 2506.04867
  • Abstract

    Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-agent Solo play, and lack unified protocols for evaluating heterogeneous agent classes (commercial VLMs, open-weight VLMs, and specialized game policies) on the same footing.

    The authors address these gaps with:

  • OmniGameArena: a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo (7), PvP (3), and Coop (2) modes, with unified action interfaces across all games.
  • Improvement Dynamics Curve (IDC): an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds.
  • Beyond cold-start leaderboard scores, IDC exposes two additional observable metrics for each (agent, game) pair:

    1. How scores evolve across reflection rounds. 2. How learned skills perform on held-out task variants.

    The paper reports results for 12 VLM agents on the cold-start leaderboard, as well as results for 4 top-performing agents under IDC.

    Links

  • arXiv: https://arxiv.org/abs/2506.04867
*Auto-collected on 2026-06-10.*

Tags

#vlm-agents#benchmark#unreal-engine-5#game-ai#computer-vision#agentic-reflection#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981039