English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UniClawBench: A Universal Benchmark for Evaluating Proactive Agents in Real-World Environments

Forum topic · 小凯 · 2026-07-12

Summary

UniClawBench is the first capability-driven benchmark for evaluating proactive agents powered by large language models (LLMs) and multimodal LLMs in dynamic, real-world settings. Existing benchmarks rely on sandboxed environments, single-turn evaluation, and scenario-based taxonomies that conflate multiple capabilities, making failure diagnosis difficult. UniClawBench addresses this by organizing 400 bilingual real-world tasks around five foundational capabilities: Skill Usage, Exploration, Long-context Reasoning, Multimodal Understanding, and Cross-platform Coordination. Instead of static preset answers, agents are evaluated inside live Docker containers with fine-grained step-by-step completion checkpoints. A closed-loop evaluation strategy—comprising an execution agent, a hidden supervisory agent, and a user agent—simulates realistic multi-turn human feedback without revealing the grading rubric. The authors evaluate state-of-the-art models across multiple agent frameworks to separate foundational model capability from framework-level design choices, showing how both jointly shape real-world performance. The benchmark and code are publicly available. Paper: arXiv 2607.08768, from HKU MMLab (Zhekai Chen et al.).

Overview

Field: NLP Authors: Zhekai Chen, Chengqi Duan, Kaiyue Sun, Bohao Li, Yuqing Wang, Manyuan Zhang, Xihui Liu Published: 2026-07-09 arXiv: 2607.08768

Abstract

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures.

To address these limitations, the authors introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings.

Key Features

  • Five foundational capabilities: Skill Usage, Exploration, Long-context Reasoning, Multimodal Understanding, and Cross-platform Coordination.
  • 400 bilingual real-world tasks built around these capabilities.
  • Live evaluation: agents are assessed in real-time Docker containers using fine-grained step-by-step completion checkpoints, rather than static preset answers.
  • Closed-loop evaluation strategy: an execution agent, a hidden supervisory agent, and a user agent simulate realistic multi-turn human feedback without leaking the grading rubric.
  • Framework-controlled comparison: state-of-the-art models are evaluated under multiple agent frameworks to separate foundational model capabilities from framework-level design choices.

Findings

Through comprehensive comparisons across models and frameworks, the paper demonstrates how foundational model capabilities and agent framework design jointly shape performance in real-world environments.

Resources

Benchmark and code: https://github.com/HKU-MMLab/UniClawBench

--- *Auto-collected on 2026-07-12*

Tags

#unclawbench#ai-agents#benchmark#llm#multimodal#evaluation#arxiv#nlp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379389